Model training method, anomaly detection method, electronic equipment and storage medium

By using the abnormal characteristics output from the first detection model to train the second detection model, the problem of lack of specific domain knowledge and insufficient training samples in the prior art is solved, and the accuracy and efficiency of abnormal detection in industrial quality inspection is improved.

CN120145293APending Publication Date: 2025-06-13BYD CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510142254.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

When existing deep learning technologies are used in industrial quality inspection, due to the lack of specific domain knowledge and insufficient training sample data, the model training effect is poor, the accuracy is low, and it is difficult to directly apply to industrial quality inspection.

Method used

By obtaining the target modal data of the target object, the second detection model is trained based on the abnormal features output by the first detection model, and using the abnormal features as prior knowledge, reducing dependence on a large amount of training data, and improving model training efficiency and accuracy.

Benefits of technology

When the amount of training sample data is small, the model training efficiency and accuracy are improved, the model's ability to identify abnormal features is enhanced, and the abnormal detection accuracy in industrial quality inspection is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145293A_ABST
    Figure CN120145293A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method, an anomaly detection method, electronic equipment and a storage medium, relates to the technical field of manufacturing quality inspection, and is used for improving the accuracy of anomaly detection in industrial quality inspection. The model training method comprises the following steps: acquiring target modal data of a target object; determining abnormal features of the target object based on the target modal data of the target object and a first detection model; training a to-be-trained second detection model based on the abnormal features of the target object and the target modal data of the target object; the second detection model is used for identifying abnormal information existing in the target object based on the target modal data of the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of manufacturing quality inspection, and particularly to a model training method, an anomaly detection method, an electronic device, and a storage medium. Background Art

[0002] In the industrial production scenario, industrial quality inspection is an important part of industrial production. Existing deep learning technologies can achieve anomaly detection of industrial production parts. However, due to the wide variety of industrial production parts, the existing general large models lack knowledge about specific fields of anomaly detection. Different types of parts require corresponding models for training, learning, and processing. However, in actual production, the available training sample data is small in quantity and cannot meet the training requirements of large-scale models, resulting in poor training effects and low accuracy of these large models, making it difficult to directly apply them to industrial quality inspection anomaly detection tasks. Summary of the Invention

[0003] Embodiments of this application provide a model training method, an anomaly detection method, an electronic device, and a storage medium, which relate to the technical field of manufacturing quality inspection and are used to improve the accuracy of anomaly detection in industrial quality inspection.

[0004] To achieve the above objective, this application adopts the following technical solutions:

[0005] In a first aspect, this application provides a model training method, which includes: obtaining target modal data of a target object; determining anomaly features of the target object based on the target modal data of the target object and a first detection model; training a second detection model to be trained based on the anomaly features of the target object and the target modal data of the target object; and the second detection model is used to identify anomaly information existing in the target object based on the target modal data of the target object.

[0006] It can be understood that the model training method provided by the embodiments of this application can train the second detection model based on the anomaly features output by the first detection model as prior knowledge. In this way, the second detection model can use the knowledge and experience of the first detection model for training without having to learn from scratch. In this way, in the case of a small amount of training sample data, the model training efficiency and accuracy are improved, and further, when the second detection model is applied to industrial quality inspection, the accuracy of anomaly detection can be improved.

[0007] In some embodiments, the second detection model includes an input layer, and the input layer includes a first input channel and a second input channel, where the first input channel is the input channel for the target modal data of the target object; and the second input channel is the input channel for the anomaly features of the target object.

[0008] In some embodiments, training the second detection model to be trained based on the abnormal features of the target object and the target modal data of the target object includes: inputting the target modal data of the target object into the second detection model to be trained to obtain the abnormal detection result of the target object; determining the value of the first loss function based on the abnormal features of the target object and the abnormal detection result; wherein, the value of the first loss function is used to reflect the difference between the abnormal features of the target object and the abnormal detection result; training the second detection model based on the value of the first loss function to obtain the trained second detection model.

[0009] In some embodiments, when the target modal data includes visual image data, the method further includes: obtaining the global features of the visual image data of the target object based on the visual image data of the target object and the global feature extraction model; training the second detection model to be trained based on the abnormal features of the target object and the target modal data of the target object, including: training the second detection model to be trained based on the visual image data of the target object, the abnormal features of the target object, and the global features of the visual image data of the target object.

[0010] In some embodiments, training the second detection model to be trained based on the visual image data of the target object, the abnormal features of the target object, and the global features of the visual image data of the target object includes: fusing the abnormal features of the target object and the global features of the visual image data of the target object to obtain the fused features; inputting the target modal data of the target object into the second detection model to be trained to obtain the abnormal detection result of the target object; determining the value of the second loss function based on the fused features and the abnormal detection result; wherein, the value of the second loss function is used to reflect the difference between the fused features and the abnormal detection result; training the second detection model based on the value of the second loss function to obtain the trained second detection model.

[0011] In some embodiments, the second detection model includes a feature fusion layer, and the feature fusion layer is used to perform feature fusion processing on the input abnormal features of the target object and the global features of the visual image data of the target object.

[0012] In some embodiments, the second detection model includes an input layer, and the input layer includes a first input channel, a second input channel, and a third input channel. Among them, the first input channel is the input channel for the target modal data of the target object; the second input channel is the input channel for the abnormal features of the target object; the third input channel is the global features of the visual image data of the target object.

[0013] In some embodiments, the first detection model is trained based on the target modal data of the target object and the target modal data annotated with the abnormal features of the target object.

[0014] In some embodiments, when the target modal data is visual image data, the method further includes: performing visualization processing on the abnormal information output by the first detection model to obtain a visualization image; and optimizing the first detection model based on the difference between the visual image data of the target object and the visualization image. In some embodiments,

[0015] In some embodiments, performing visualization processing on the abnormal information output by the first detection model to obtain a visualization image includes: performing visualization processing on the abnormal information by using the gradient-weighted class activation mapping method to obtain a visualization image.

[0016] In some embodiments, the visualization image is a heat map.

[0017] In some embodiments, the abnormal features include at least one of the following: abnormal category, abnormal region coordinates, confidence score.

[0018] In some embodiments, the scale of the second detection model is larger than that of the first detection model.

[0019] In some embodiments, the target object includes automotive parts.

[0020] In some embodiments, the target modal data includes at least one of the following: visual image data, sound data, vibration data.

[0021] In a second aspect, the present application provides an abnormal detection method, the method including: obtaining target modal data of a target object to be detected; inputting the target modal data into a second detection model to obtain abnormal information of the target object; where the second detection model is trained by using the model training method provided in the first aspect and any one of its embodiments.

[0022] In some embodiments, the target object includes automotive parts.

[0023] In some embodiments, the target modal data includes at least one of the following: visual image data, sound data, vibration data.

[0024] In a third aspect, the present application provides a model training device, the model training device including: a communication module, a processing module, and a training module.

[0025] The communication module is configured to obtain target modal data of a target object;

[0026] The processing module is configured to determine abnormal features of the target object based on the target modal data of the target object and a first detection model;

[0027] A training module, configured to train a second detection model to be trained based on the abnormal features of a target object and the target modal data of the target object; the second detection model is configured to identify the abnormal information existing in the target object based on the target modal data of the target object.

[0028] In some embodiments, the second detection model includes an input layer, and the input layer includes a first input channel and a second input channel, where the first input channel is the input channel for the target modal data of the target object; the second input channel is the input channel for the abnormal features of the target object.

[0029] In some embodiments, the training module is specifically configured to input the target modal data of the target object into the second detection model to be trained to obtain the abnormal detection result of the target object; determine the value of a first loss function based on the abnormal features of the target object and the abnormal detection result; where the value of the first loss function is used to reflect the difference between the abnormal features of the target object and the abnormal detection result; train the second detection model based on the value of the first loss function to obtain the trained second detection model.

[0030] In some embodiments, when the target modal data includes visual image data, the processing module is further configured to obtain the global features of the visual image data of the target object based on the visual image data of the target object and a global feature extraction model; the training module is specifically configured to train the second detection model to be trained based on the visual image data of the target object, the abnormal features of the target object, and the global features of the visual image data of the target object.

[0031] In some embodiments, the training module is specifically configured to fuse the abnormal features of the target object and the global features of the visual image data of the target object to obtain fused features; input the target modal data of the target object into the second detection model to be trained to obtain the abnormal detection result of the target object; determine the value of a second loss function based on the fused features and the abnormal detection result; where the value of the second loss function is used to reflect the difference between the fused features and the abnormal detection result; train the second detection model based on the value of the second loss function to obtain the trained second detection model.

[0032] In some embodiments, the second detection model includes a feature fusion layer, and the feature fusion layer is configured to perform feature fusion processing on the input abnormal features of the target object and the global features of the visual image data of the target object.

[0033] In some embodiments, the second detection model includes an input layer, and the input layer includes a first input channel, a second input channel, and a third input channel. Among them, the first input channel is the input channel for the target modal data of the target object; the second input channel is the input channel for the abnormal features of the target object; the third input channel is the global feature of the visual image data of the target object.

[0034] In some embodiments, the first detection model is trained based on the target modal data of the target object and the target modal data annotated with the abnormal features of the target object.

[0035] In some embodiments, when the target modal data is visual image data, the processing module is further configured to perform visualization processing on the abnormal information output by the first detection model to obtain a visualization image; and optimize the first detection model based on the difference between the visual image data of the target object and the visualization image.

[0036] In some embodiments, the processing module is specifically configured to perform visualization processing on the abnormal information by using the gradient-weighted class activation mapping method to obtain a visualization image.

[0037] In some embodiments, the visualization image is a heat map.

[0038] In some embodiments, the abnormal features include at least one of the following: abnormal category, abnormal region coordinates, confidence score.

[0039] In some embodiments, the scale of the second detection model is larger than that of the first detection model.

[0040] In some embodiments, the target object includes automotive parts.

[0041] In some embodiments, the target modal data includes at least one of the following: visual image data, sound data, vibration data.

[0042] In a fourth aspect, the present application provides an electronic device, which includes: a processor and a memory; the memory stores instructions executable by the processor; when the processor is configured to execute the instructions, the electronic device implements the method of the first aspect above.

[0043] In a fifth aspect, the present application provides a computer-readable storage medium, which includes: computer software instructions; when the computer software instructions run on an electronic device, the electronic device implements the method of the first aspect above.

[0044] In a sixth aspect, the present application provides a computer program product, which includes a computer program; when the computer program runs on an electronic device, the electronic device implements the method of the first aspect above.

[0045] For the beneficial effects of the second to sixth aspects above, reference may be made to the corresponding descriptions of the first aspect, and details are not repeated herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0047] Figure 1 Schematic diagram of the composition of an anomaly detection system provided by an embodiment of the present application;

[0048] Figure 2 Schematic diagram of the process of a model training method provided by an embodiment of the present application;

[0049] Figure 3 A set of product images with anomaly features provided by an embodiment of the present application;

[0050] Figure 4 Schematic diagram of the process of another model training method provided by an embodiment of the present application;

[0051] Figure 5 Schematic diagram of the process of yet another model training method provided by an embodiment of the present application;

[0052] Figure 6 Schematic diagram of the process of yet another model training method provided by an embodiment of the present application;

[0053] Figure 7 Schematic diagram of the process of yet another model training method provided by an embodiment of the present application;

[0054] Figure 8 Schematic diagram of the process of yet another model training method provided by an embodiment of the present application;

[0055] Figure 9 Visual image of a target object provided by an embodiment of the present application;

[0056] Figure 10 Visualization image of a target object provided by an embodiment of the present application;

[0057] Figure 11 Schematic diagram of the process of yet another model training method provided by an embodiment of the present application;

[0058] Figure 12 Schematic diagram of the process of yet another model training method provided by an embodiment of the present application;

[0059] Figure 13 Schematic flowchart of another anomaly detection method provided by an embodiment of the present application;

[0060] Figure 14 Schematic structural diagram of a model training device provided by an embodiment of the present application;

[0061] Figure 15 Schematic structural diagram of an electronic device provided by the present application.

[0062] Reference numerals: data acquisition device 101, training device 102, detection device 103. Detailed implementation manners

[0063] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0064] In the description of the present application, it should be noted that unless otherwise clearly defined and limited, the terms "installed", "connected", "connected", and "communicated" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection. It may be directly connected, or indirectly connected through an intermediate medium, and it may be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific situations.

[0065] In the embodiments of the present application, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, article or device. Without further limitation, the element defined by the statement "including one..." does not exclude the existence of other identical elements in the process, article or device including the element.

[0066] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0067] In the description of this specification, specific features, structures, materials, or characteristics may be combined in a suitable manner in any one or more embodiments or examples.

[0068] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by this application, and all should be covered by the protection scope of this application.

[0069] First, some professional data designed for this application will be introduced.

[0070] BLIP-2: It is an efficient vision-language pre-training model that guides vision-language pre-training by leveraging off-the-shelf frozen pre-trained image encoders and large language models. The core of BLIP-2 is a lightweight query Transformer (Q-Former), which is pre-trained in two stages: The first stage: guiding vision-language representation learning from the frozen image encoder, and the Q-Former learns the visual representations most relevant to the text. The second stage: guiding vision-to-language generation learning from the frozen large language model. By connecting the output of the Q-Former to the frozen large language model and training the Q-Former so that the visual representations it outputs can be interpreted by the large language model.

[0071] Yolo-v8: It is a one-stage object detection algorithm that directly performs classification and localization on a single image and can be used for tasks such as image classification, object detection, and instance segmentation.

[0072] Cross Attention: It is an attention mechanism used in deep learning to process two different input sequences. It is particularly important in Transformer models, especially in sequence-to-sequence tasks such as machine translation, question answering systems, image captioning, etc.

[0073] Gradient-weighted Class Activation Mapping (Grad-CAM) is a technique for visualizing the decision-making process of convolutional neural networks. It highlights the regions in the input image that are important for the model to predict a specific class by generating a heatmap. Since Grad-CAM does not require changes to the model architecture or retraining, it can be widely applied to various convolutional neural network architectures.

[0074] As described in the background art, in the production scenarios of the traditional automotive industry, due to reasons such as low technological maturity, long time-consuming for system operation to get started, and untimely production dynamic monitoring, it often results in excessive manual input, serious material loss, and low management efficiency, leading to high factory production costs and low efficiency. Therefore, industrial quality inspection based on the intelligent manufacturing scenario is an indispensable part of industrial production. Existing AI deep learning technologies can achieve anomaly detection of components and provide anomaly scores at the same time, but manual threshold setting is required to distinguish normal and abnormal samples, and this operation limits the actual application scenarios of AI algorithms. In addition, there are a wide variety of components in automobile factories, and corresponding AI models are required for training, learning, and processing different types of components. Moreover, a large amount of data is required during the training process of the models. These objective conditions result in relatively high production, maintenance, and update costs of the AI system.

[0075] Existing large model technologies can be widely applied to various visual tasks such as image description, visual understanding, and visual reasoning through the combination of deep learning and natural language processing technologies, showing excellent perception capabilities. However, due to the lack of knowledge in specific fields related to anomaly detection and the weak understanding of local details of images in existing general large models, and in actual production, the available number of training sample data is small and cannot meet the training requirements of large-scale models, resulting in poor training effects and low accuracy of these large models, making it difficult to directly apply them to industrial quality inspection anomaly detection tasks.

[0076] To address the above technical problems, this application proposes a model training method. The idea of this method is as follows: obtaining the target modal data of the target object; determining the anomaly features of the target object based on the target modal data of the target object and the first detection model; training the second detection model to be trained based on the anomaly features of the target object and the target modal data of the target object; the second detection model is used to identify the anomaly information existing in the target object based on the target modal data of the target object. The model training method provided by the embodiments of this application can train the second detection model based on the anomaly features output by the first detection model as prior knowledge. In this way, the second detection model can be trained using the knowledge and experience of the first detection model without having to start learning from scratch. In this way, when the amount of training sample data is small, the model training efficiency and accuracy can be improved. Furthermore, when the second detection model is applied to industrial quality inspection, the accuracy of anomaly detection can be improved.

[0077] Figure 1 This is a schematic diagram of the composition of an anomaly detection system provided by the embodiments of this application. As Figure 1 described, it includes: a data acquisition device 101, a training device 102, and a detection device 103. The data acquisition device 101 is communicatively connected to the training device 102 and the detection device 103 respectively, and the training device 102 and the detection device 103 are communicatively connected.

[0078] The data acquisition device 101 is used to obtain the target modal data of the target object.

[0079] In some embodiments, the target modal data includes at least one of the following: visual image data, sound data, vibration data.

[0080] In some embodiments, the data acquisition device 101 includes at least one of the following: a high-definition photography device for collecting visual image data, a vibration sensor data for extracting the vibration feature information of the target object, and an acoustic wave sensor for extracting the sound feature information (sound data) of the target object. The specific form of the data acquisition device 101 in this application is not limited, Figure 1 and is shown in the form of a photography device.

[0081] The training device 102 is used to collect the target modal data of the target object and perform processing, including data cleaning, data preprocessing, and data annotation. The processed target modal data is used to train the anomaly detection model to be trained.

[0082] Exemplarily, assuming that the screw is the target object, the collected visual image data contains images of various screws and their potential defects, such as damaged threads, cracked heads, and inconsistent dimensions.

[0083] Exemplarily, the training device 102 can be a server cluster composed of multiple servers, or a single server, or a computer, or a processor or processing chip in the server or computer, etc. The specific device form of the training device 102 in the embodiments of this application is not limited.

[0084] The training device 102 is used to determine the anomaly features of the target object based on the target modal data of the target object and the first detection model.

[0085] In some embodiments, the training device 102 is further used to perform visualization processing on the anomaly information output by the first detection model to obtain a visualization image; and optimize the first detection model based on the difference between the visual image data of the target object and the visualization image.

[0086] In some embodiments, the visualization image is a heat map.

[0087] In some embodiments, when the target modal data includes visual image data, the method further includes: obtaining the global features of the visual image data of the target object based on the visual image data of the target object and the global feature extraction model.

[0088] Exemplarily, the training device 102 may be a server cluster composed of multiple servers, or a single server, or a computer, or a processor or processing chip in a server or computer, etc. The embodiments of the present application do not limit the specific device form of the training device 102.

[0089] The training device 102 is used to train a second detection model to be trained based on the abnormal features of the target object and the target modal data of the target object; the second detection model is used to identify the abnormal information existing in the target object based on the target modal data of the target object.

[0090] In some embodiments, the training device 102 is further configured to input the target modal data of the target object into the second detection model to be trained, obtain the abnormal detection result of the target object; determine the value of a first loss function based on the abnormal features of the target object and the abnormal detection result; wherein, the value of the first loss function is used to reflect the difference between the abnormal features of the target object and the abnormal detection result; train the second detection model based on the value of the first loss function to obtain the trained second detection model.

[0091] In some embodiments, the training device 102 is further configured to train the second detection model to be trained based on the visual image data of the target object, the abnormal features of the target object, and the global features of the visual image data of the target object.

[0092] Exemplarily, the training device 102 may be a server cluster composed of multiple servers, or a single server, or a computer, or a processor or processing chip in a server or computer, etc. The embodiments of the present application do not limit the specific device form of the training device 102.

[0093] The detection device 103, in which a trained second detection model is deployed, can determine the abnormal information of the target object based on the target modal data of the target object collected in real time and the second detection model.

[0094] Exemplarily, the detection device 103 may be a server cluster composed of multiple servers, or a single server, or a computer, or a processor or processing chip in a server or computer, etc. The embodiments of the present application do not limit the specific device form of the detection device 103.

[0095] It should be noted that Figure 1 is only an exemplary framework diagram, Figure 1 the number of devices included therein, the names of each device are not limited, and in addition to Figure 1 the devices shown, other devices may also be included, and the embodiments of the present application do not limit this.

[0096] It should be noted that the application scenarios of the embodiments of the present disclosure are not limited. The system architecture and business scenarios described in the embodiments of the present disclosure are for more clearly explaining the technical solutions of the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Those of ordinary skill in the art know that with the evolution of the network architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present disclosure are equally applicable to similar technical problems.

[0097] The following specifically introduces the model training method provided by this application in conjunction with the accompanying drawings.

[0098] See Figure 2 , which is a schematic flowchart of a model training method provided by this application, and is applied to a training device as shown in Figure 1 As shown in Figure 2 As shown, the method includes:

[0099] S101. Obtain the target modal data of the target object.

[0100] In some embodiments, the target object includes products in the manufacturing field, such as auto parts.

[0101] In some embodiments, the target modal data includes at least one of the following: visual image data, sound data, vibration data.

[0102] S102. Determine the abnormal features of the target object based on the target modal data of the target object and the first detection model.

[0103] Among them, the first detection model is a feature extraction model with a relatively small total number of parameters, and is used to extract the abnormal features in the target modal data based on the target modal data of the target object. Exemplarily, the first detection model can select Yolo-v8, and the training process of using Yolo-v8 as the first detection model is introduced in the subsequent steps S1-S4, which will not be elaborated here.

[0104] In some embodiments, the abnormal features include at least one of the following: abnormal category, abnormal area coordinates, confidence score.

[0105] Exemplarily, for the manufacturing quality inspection scenario, common abnormal categories include at least one of the following: burrs on the product surface, non-conforming product dimensions, product defects, etc. As shown in Figure 3 , a1 is an image of burrs on the product surface, and a2 is an image of a defective product.

[0106] Exemplarily, in the case where the target modal data includes visual image data, the abnormal region coordinates represent the coordinates of the center point of the region where the abnormal feature is located in the visual image relative to the visual image, as well as the width and height of the region where the abnormal feature is located.

[0107] Exemplarily, the confidence score represents the detection reliability of the first detection model for the abnormal features of the target object.

[0108] S103. Train the second detection model to be trained based on the abnormal features of the target object and the target modal data of the target object.

[0109] In some embodiments, the scale of the second detection model is larger than that of the first detection model. Exemplarily, the second detection model adopts a more complex network structure and a larger number of parameters than the first detection model. At the same time, the number of samples for training the second detection model is usually larger than the number of samples for training the first detection model.

[0110] In some embodiments, the second detection model is used to identify the abnormal information existing in the target object based on the target modal data of the target object.

[0111] In some embodiments, the second detection model includes an input layer, and the input layer includes a first input channel and a second input channel. Among them, the first input channel is the input channel for the target modal data of the target object; the second input channel is the input channel for the abnormal features of the target object.

[0112] Exemplarily, taking BLIP-2 as the second detection model and visual image data as the modal data of the target object, the above-mentioned training of the second detection model to be trained based on the abnormal features of the target object and the modal data of the target object has the following process: First, through three pre-training tasks, namely the image-text contrastive learning task, the image-text matching task, and the image-based text generation task, help the Q-Former in the BLIP-2 model learn visual and language representations. Then, connect the output of the Q-Former module with the large language model with frozen parameters, and train the Q-Former module to generate text consistent with the image content. Among them, for the large language model based on the pure decoder architecture (such as OPT), the language modeling objective function is used for training, and the large language model with frozen parameters generates text according to the visual representation provided by the Q-Former. For the large language model based on the encoder-decoder architecture (such as FlanT5), the text is divided into a prefix and a suffix. The prefix and the output of the Q-Former are used as the input of the encoder, and the decoder outputs the suffix.

[0113] In some embodiments, the first detection model is trained based on the target modal data of the target object and the target modal data annotated with the abnormal features of the target object.

[0114] Exemplarily, as Figure 4 shown, taking Yolo-v8 as the first detection model and visual image data as the modal data of the target object as an example, the training process of the first detection model is as follows:

[0115] S1. Create a configuration file.

[0116] In some embodiments, creating a configuration file includes creating and editing a data.yaml file (usually used to store the configuration information of the dataset), defining the path of the dataset, class names, training / validation / test set division, etc. At the same time, according to project requirements, hyperparameters, learning rate, number of iterations, and image size can be adjusted before training the model.

[0117] S2. Model training.

[0118] In some embodiments, taking visual image data as samples and visual image data annotated with abnormal features as sample labels, the Yolo-v8 model is trained.

[0119] Exemplarily, use the command-line tool to start the training process, and the command is as follows: "python train.py --data data.yaml --weights yolov8s.pt --img 640 --epochs 100". Among them, "-img 640" specifies the size of the input image, and "--epochs 100" specifies the number of training rounds.

[0120] S3. Judge whether the mAP and recall metrics meet the expectations.

[0121] In some embodiments, during and after the training process of the model, the performance of the model is evaluated using the validation set. Judge whether the mAP and recall metrics meet the expectations. If they do not meet the expectations, continue training (execute step a2) and adjust the model (such as adjusting the learning rate, adjusting the model architecture, etc.). If they meet the expectations, continue to execute step a4.

[0122] S4. Deploy the model and generate a feature map.

[0123] In some embodiments, export the trained model, write feature extraction code, generate the feature map of each training sample (i.e., visual image) and save it for subsequent training.

[0124] In some embodiments, such asFigure 5 As shown, the above S103 can be specifically implemented as follows:

[0125] S201. Input the target modal data of the target object into the second detection model to be trained, and obtain the anomaly detection result of the target object.

[0126] S202. Determine the value of the first loss function based on the anomaly features and the anomaly detection result of the target object.

[0127] Among them, the value of the first loss function is used to reflect the difference between the anomaly features of the target object and the anomaly detection result.

[0128] S203. Train the second detection model based on the value of the first loss function to obtain the trained second detection model.

[0129] Exemplarily, the cross-entropy loss function is selected as the first loss function. The cross-entropy loss function optimizes the model by measuring the difference between the predicted probability distribution of the model and the true label distribution. The formula of the cross-entropy loss function is as follows:

[0130]

[0131] Among them, C is the number of categories, y i is the true label (using one-hot encoding, 1 represents the correct category, and 0 represents other categories), is the probability of the i-th category predicted by the model (usually the probability value output by the softmax function).

[0132] It can be understood that the cross-entropy loss function can handle classification problems (in this application, it refers to distinguishing normal parts from abnormal parts). By measuring the difference between the predicted distribution and the true distribution, it helps the model adjust the weights during the training process to minimize the prediction error and improve the accuracy of the model.

[0133] In some embodiments, the target modal data includes visual image data. In this case, as Figure 6 shown, before step S103, the model training method provided in this application further includes:

[0134] S301. Based on the visual image data of the target object and the global feature extraction model, obtain the global features of the visual image data of the target object.

[0135] Exemplarily, the global feature extraction model can be a pre-trained Vision Transformer (ViT) model.

[0136] Based on S301, the above S103 can be specifically implemented as follows:

[0137] S302. Train the second detection model to be trained based on the visual image data of the target object, the abnormal features of the target object, and the global features of the visual image data of the target object.

[0138] In some embodiments, the second detection model includes an input layer, and the input layer includes a first input channel, a second input channel, and a third input channel. Among them, the first input channel is the input channel for the target modal data of the target object; the second input channel is the input channel for the abnormal features of the target object; the third input channel is the global feature of the visual image data of the target object.

[0139] It can be understood that by fusing the global features and the abnormal features, the model's attention and understanding of the abnormal features can be enhanced, enabling the model to not only focus on the overall structure of the data but also be more sensitive to capturing abnormal features, helping the model better identify abnormal parts when processing complex data and improving the accuracy of the model in identifying abnormal features.

[0140] Furthermore, as Figure 7 shown, step S302 can be specifically implemented as:

[0141] S401. Fuse the abnormal features of the target object and the global features of the visual image data of the target object to obtain fused features.

[0142] In some embodiments, the second detection model includes a feature fusion layer, and the feature fusion layer is used to perform feature fusion processing on the input abnormal features of the target object and the global features of the visual image data of the target object.

[0143] Exemplarily, the feature fusion layer implements feature fusion processing on the abnormal features of the target object and the global features of the visual image data of the target object through a cross-attention mechanism (Cross Attention). The fusion process is as follows: First, input the abnormal features of the target object and the global features of the visual image data of the target object into the fusion module. Inside the module, through the channel attention cross mechanism, the importance of the channels of each modal feature is evaluated and weighted to highlight the key channel information; then, in the spatial dimension, the spatial attention fusion mechanism is used to enable the model to focus on the abnormal region and enhance the perception of abnormal features. At the same time, a learnable weight parameter α is introduced to dynamically adjust the weight between the abnormal features and the global features, and α is optimized through backpropagation to adapt to the feature fusion requirements in different situations. Finally, the weighted features are fused to obtain fused features.

[0144] In some embodiments, fusing the abnormal features of the target object and the global features of the visual image data of the target object to obtain fused features further includes: fusing the abnormal feature vector of the target object and the global features of the visual image data of the target object to obtain fused features.

[0145] Exemplarily, the abnormal features include an abnormal category, abnormal region coordinates, and a confidence score. The process of converting the abnormal features into abnormal feature quantities is as follows: encoding the category identifier of the abnormal category into a vector form, such as using one-hot encoding. For visual image data, mapping the abnormal region coordinate information to the pixel space of the image to form a feature map or mask. Incorporating the confidence score directly into the feature vector. Integrating multiple abnormal features into a multi-dimensional feature vector (i.e., the abnormal feature vector) to achieve normalization or standardization processing of the abnormal features.

[0146] S402. Input the target modal data of the target object into the second detection model to be trained, and obtain the abnormal detection result of the target object.

[0147] S403. Based on the fused features and the abnormal detection result, determine the value of the second loss function.

[0148] Wherein, the value of the second loss function is used to reflect the difference between the fused features and the abnormal detection result.

[0149] S404. Train the second detection model based on the value of the second loss function to obtain the trained second detection model.

[0150] In some embodiments, in the scenario of component detection, the negative sample data in the training samples is small, and the number of positive and negative samples is unbalanced. During the process of training the second detection model using the fused features, there may be a problem of model overfitting. Therefore, a second loss function is introduced to improve the problem of unbalanced positive and negative samples.

[0151] Exemplarily, select the Focal Loss function as the second loss function. It uses a modulating factor and a balancing parameter to adjust the weights of positive and negative samples, and solves the problem of unbalanced positive and negative samples at one time. The formula of the Focal Loss function is as follows:

[0152] FL(p t )=-α t (1-p t ) γ log(p t ),

[0153] Wherein, p t is the prediction probability of the model for the sample (for a positive sample, p t is the probability of predicting as positive; for a negative sample, pt is the complement of the probability predicted as negative), α t is a parameter for adjusting the weights of positive and negative samples, and γ is a parameter for adjusting the sample weights (also known as the focusing parameter).

[0154] It can be understood that based on the combined loss function of Cross-Entropy Loss and Focal Loss, this application enables the model to simultaneously consider the joint optimization objective of the image and abnormal features, increases the penalty term for the prediction accuracy of abnormal regions, and ensures that while the model learns the knowledge related to the learning task, it can also strengthen the recognition and understanding of abnormal situations.

[0155] In some embodiments, when the target modal data is visual image data, such as Figure 8 shown, the model training method provided by this application further includes:

[0156] S501. Visualize the abnormal information output by the first detection model to obtain a visualization image.

[0157] In some embodiments, the abnormal information is visualized by the Gradient-Weighted Class Activation Mapping method (Grad-CAM) to obtain a visualization image.

[0158] In some embodiments, the visualization image is a heatmap.

[0159] S502. Optimize the first detection model based on the difference between the visual image data of the target object and the visualization image.

[0160] It can be understood that the visualization image (such as a heatmap) can intuitively display the location and degree of product abnormalities. By comparing the visualization image with the original visual image data, the deviation of model detection can be quickly and intuitively discovered, providing a clear direction for model optimization.

[0161] Exemplarily, the process of visualizing the abnormal information output by the first detection model through Grad-CAM is as follows:

[0162] b1. Obtain the gradient with respect to the last convolutional layer of the first detection model.

[0163] b2. Obtain the class activation map.

[0164] In some embodiments, the importance weight of each position is obtained by multiplying the activation of the last convolutional layer by the corresponding gradient, and then global average pooling is performed on the weighted activation to obtain the class activation map.

[0165] b3. Generate a heatmap.

[0166] In some embodiments, the class activation map is normalized, and using a color map (such as the jet color map), the class activation map is converted into a heatmap.

[0167] b4. Implement visualization.

[0168] In some embodiments, the heatmap is superimposed on the original image, and the superimposed image is output, thus completing the visualization of the abnormal information.

[0169] In some embodiments, based on the difference between the visual image data of the target object and the visualization image, the parameters and architecture of the first detection model are adjusted according to historical experience to optimize the first detection model.

[0170] Exemplarily, as Figure 9 shown is a visual image of a target object provided by an embodiment of the present application, Figure 10 and shown is a visualization image of a target object provided by an embodiment of the present application.

[0171] Based on the above model training method, as Figure 11 shown, the present application further provides an abnormal detection method, as shown in the figure, including:

[0172] S601. Obtain the target modal data of the target object to be detected.

[0173] S602. Input the target modal data into the second detection model to obtain the abnormal information of the target object.

[0174] Among them, the second detection model is trained based on the above model training method.

[0175] In some embodiments, the target object includes automotive parts.

[0176] In some embodiments, the target modal data includes at least one of the following: visual image data, sound data, vibration data.

[0177] The model training method and the abnormal detection method of the present application are introduced below in conjunction with the process schematic diagram.

[0178] Figure 12 shown is a process schematic diagram of a model training method provided by an embodiment of the present application, as Figure 12As shown, assuming that the target modal data of the target object is visual image data, the visual image data is input into the global feature extraction model and the first detection model respectively to extract the global feature and the abnormal feature of the visual image data. Then, the global feature and the abnormal feature are sent to the feature fusion layer respectively, and the feature fusion layer fuses the global feature and the abnormal feature to obtain a fused feature. Then, based on the fused feature and the visual image data, the second detection model is trained to obtain the trained second detection model.

[0179] Figure 13 It is a schematic flowchart of an anomaly detection method provided by an embodiment of the present application. As Figure 13 shown, the trained second detection model is connected to the large language model, and the large language model is used to implement the interaction with the user (quality inspector). The detection process is as follows: The user sends the target modal data of the target object and the user detection instruction to the large language model. The large language model generates a detection instruction that can be recognized by the second detection model based on the text detection instruction, and sends the target modal data and the detection instruction to the second detection model. After the second detection model completes the detection task according to the detection instruction, it returns the anomaly information to the large language model. Finally, the large language model generates a description of the anomaly information (including the anomaly category) based on the anomaly information and returns it to the user.

[0180] The model training method provided by the present application first obtains the target modal data of the target object, such as visual image data, sound data, or vibration data in product detection in the manufacturing field. Then, based on the target modal data and the first detection model, the abnormal features of the target object are determined, including the abnormal category, the abnormal area coordinates, and the confidence score, etc. Next, based on the abnormal features and the target modal data of the target object, a larger-scale and more complex second detection model is trained. During the training process, the anomaly detection result can be obtained by inputting the target modal data, and the value of the loss function is determined based on the abnormal features and the detection result, so as to optimize the training of the model. If the target modal data is visual image data, the global feature can also be obtained through the global feature extraction model, and the global feature and the abnormal feature are fused and then used for the training of the second detection model. In addition, the anomaly information output by the first detection model can be visualized, and the first detection model can be optimized according to the difference between the visualized image and the original visual image data. Compared with the prior art, this method can use a small-scale model to extract sample data first in a scenario with a small number of samples to enhance the features. When using a small number of samples to train a large-scale model, the large-scale model can efficiently learn the features between the features and the samples, improve the accuracy of anomaly recognition, and solve the problem that the training effect of the large-scale model is poor and the accuracy is low when training the large-scale model with a small number of samples.

[0181] It can be seen that the above mainly introduces the solution provided by the embodiments of the present application from the perspective of methods. To implement the above functions, the embodiments of the present application provide the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, in combination with the modules and algorithm steps of each example described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0182] The embodiments of the present application can divide the function modules of the model training device according to the above method examples. For example, each function module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software function modules. Optionally, the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there can be other division methods in actual implementation.

[0183] Figure 14 It is a schematic structural diagram of a model training device provided by an embodiment of the present application, which can implement the model training method provided by the above method embodiment. As Figure 14 shown, the model training device 500 includes: a communication module 501, a processing module 502, and a training module 503.

[0184] The communication module 501 is used to obtain the target modal data of the target object;

[0185] The processing module 502 is used to determine the abnormal features of the target object based on the target modal data of the target object and the first detection model;

[0186] The training module 503 is used to train the second detection model to be trained based on the abnormal features of the target object and the target modal data of the target object; the second detection model is used to identify the abnormal information existing in the target object based on the target modal data of the target object.

[0187] In some embodiments, the second detection model includes an input layer, and the input layer includes a first input channel and a second input channel, where the first input channel is the input channel of the target modal data of the target object; the second input channel is the input channel of the abnormal features of the target object.

[0188] In some embodiments, the training module 503 is specifically configured to input the target modal data of the target object into the second detection model to be trained, and obtain the anomaly detection result of the target object; determine the value of the first loss function based on the anomaly feature and the anomaly detection result of the target object; wherein, the value of the first loss function is used to reflect the difference between the anomaly feature of the target object and the anomaly detection result; train the second detection model based on the value of the first loss function to obtain the trained second detection model.

[0189] In some embodiments, when the target modal data includes visual image data, the processing module 502 is further configured to obtain the global feature of the visual image data of the target object based on the visual image data of the target object and the global feature extraction model; the training module 503 is specifically configured to train the second detection model to be trained based on the visual image data of the target object, the anomaly feature of the target object, and the global feature of the visual image data of the target object.

[0190] In some embodiments, the training module 503 is specifically configured to fuse the anomaly feature of the target object and the global feature of the visual image data of the target object to obtain a fused feature; input the target modal data of the target object into the second detection model to be trained, and obtain the anomaly detection result of the target object; determine the value of the second loss function based on the fused feature and the anomaly detection result; wherein, the value of the second loss function is used to reflect the difference between the fused feature and the anomaly detection result; train the second detection model based on the value of the second loss function to obtain the trained second detection model.

[0191] In some embodiments, the second detection model includes a feature fusion layer, and the feature fusion layer is configured to perform feature fusion processing on the input anomaly feature of the target object and the global feature of the visual image data of the target object.

[0192] In some embodiments, the second detection model includes an input layer, and the input layer includes a first input channel, a second input channel, and a third input channel. Among them, the first input channel is the input channel for the target modal data of the target object; the second input channel is the input channel for the anomaly feature of the target object; the third input channel is the global feature of the visual image data of the target object.

[0193] In some embodiments, the first detection model is trained based on the target modal data of the target object and the target modal data annotated with the anomaly feature of the target object.

[0194] In some embodiments, when the target modal data is visual image data, the processing module 502 is further configured to perform visualization processing on the abnormal information output by the first detection model to obtain a visualization image; and optimize the first detection model based on the difference between the visual image data of the target object and the visualization image.

[0195] In some embodiments, the processing module 502 is specifically configured to perform visualization processing on the abnormal information by using the gradient-weighted class activation mapping method to obtain a visualization image.

[0196] In some embodiments, the visualization image is a heat map.

[0197] In some embodiments, the abnormal features include at least one of the following: abnormal category, abnormal region coordinates, confidence score.

[0198] In some embodiments, the scale of the second detection model is larger than that of the first detection model.

[0199] In some embodiments, the target object includes automotive parts.

[0200] In some embodiments, the target modal data includes at least one of the following: visual image data, sound data, vibration data.

[0201] When the functions of the above integrated modules are implemented in the form of hardware, an exemplary structural diagram of the electronic device involved in the above embodiments is provided in the embodiments of the present invention. As Figure 15 shown, the electronic device 900 includes: a processor 902, a communication interface 903, and a bus 904. Optionally, the electronic device 900 may further include a memory 901.

[0202] The processor 902 may be configured to implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of the present application. The processor 902 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It may implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of the present application. The processor 902 may also be a combination of computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0203] The communication interface 903 is configured to connect to other devices through a communication network. The communication network may be an Ethernet, a wireless access network, a wireless local area network (WLAN), etc.

[0204] The memory 901 can be a read-only memory (ROM), or other types of static storage devices that can store static information and instructions, a random access memory (RAM), or other types of dynamic storage devices that can store information and instructions. It can also be an electrically erasable programmable read-only memory (EEPROM), a magnetic disk storage medium, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0205] As a possible implementation, the memory 901 can exist independently of the processor 902. The memory 901 can be connected to the processor 902 through the bus 904 for storing instructions or program code. When the processor 902 calls and executes the instructions or program code stored in the memory 901, the model training method provided by the embodiments of the present invention can be implemented.

[0206] In another possible implementation, the memory 901 can also be integrated with the processor 902.

[0207] The bus 904 can be an extended industry standard architecture (EISA) bus, etc. The bus 904 can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, Figure 15 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0208] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the service call device is divided into different functional modules to complete all or part of the functions described above.

[0209] The embodiments of the present application also provide a computer-readable storage medium. All or part of the processes in the above method embodiments can be instructed by computer program instructions to complete the relevant hardware. The program can be stored in the above computer-readable storage medium. When the computer program instructions are executed on the computer, the computer executes the model training method described in any one of the above embodiments.

[0210] Exemplarily, the above computer-readable storage medium may include, but is not limited to: magnetic storage devices (such as hard disks, floppy disks, or magnetic tapes, etc.), optical discs (such as Compact Discs (CDs), Digital Versatile Discs (DVDs), etc.), smart cards, and flash memory devices (such as Erasable Programmable Read-Only Memories (EPROMs), cards, sticks, or key drives, etc.). The various computer-readable storage media described in this disclosure may represent one or more devices and / or other machine-readable storage media for storing information. The term "machine-readable storage medium" may include, but is not limited to, wireless channels and various other media capable of storing, containing, and / or carrying instructions and / or data).

[0211] An embodiment of the present application further provides a computer program product. The computer program product includes a computer program. When the computer program product runs on a computer, the computer is caused to execute any one of the model training methods provided in the above embodiments.

[0212] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A model training method, characterized in that: The method comprises: Obtain target modality data of the target object; Determining abnormal features of the target object based on the target modality data of the target object and a first detection model; Based on the abnormal features of the target object and the target modal data of the target object, a second detection model to be trained is trained; the second detection model is used to identify abnormal information existing in the target object based on the target modal data of the target object.

2. The method according to claim 1, characterized in that The second detection model includes an input layer, and the input layer includes a first input channel and a second input channel, wherein the first input channel is an input channel for target modal data of the target object; and the second input channel is an input channel for abnormal features of the target object.

3. The method according to claim 1, characterized in that The training of the second detection model to be trained based on the abnormal feature of the target object and the target modality data of the target object includes: Inputting the target modality data of the target object into the second detection model to be trained to obtain an abnormality detection result of the target object; Based on the abnormal characteristics of the target object and the abnormal detection result, determining a value of a first loss function; wherein the value of the first loss function is used to reflect the difference between the abnormal characteristics of the target object and the abnormal detection result; The second detection model is trained based on the value of the first loss function to obtain the trained second detection model.

4. The method according to claim 1, characterized in that: In the case where the target modality data includes visual image data, the method further includes: Based on the visual image data of the target object and a global feature extraction model, obtaining global features of the visual image data of the target object; The training of the second detection model to be trained based on the abnormal feature of the target object and the target modality data of the target object includes: The second detection model to be trained is trained based on the visual image data of the target object, the abnormal features of the target object and the global features of the visual image data of the target object.

5. The method according to claim 4, characterized in that The training of the second detection model to be trained based on the visual image data of the target object, the abnormal features of the target object and the global features of the visual image data of the target object comprises: Fusing the abnormal features of the target object with the global features of the visual image data of the target object to obtain a fused feature; Inputting the target modality data of the target object into the second detection model to be trained to obtain an abnormality detection result of the target object; Based on the fused feature and the anomaly detection result, determining a value of a second loss function; wherein the value of the second loss function is used to reflect the difference between the fused feature and the anomaly detection result; The second detection model is trained based on the value of the second loss function to obtain the trained second detection model.

6. The method according to claim 5, characterized in that The second detection model includes a feature fusion layer, and the feature fusion layer is used to perform feature fusion processing on the input abnormal features of the target object and the global features of the visual image data of the target object.

7. The method according to claim 4, characterized in that The second detection model includes an input layer, which includes a first input channel, a second input channel and a third input channel, wherein the first input channel is an input channel for target modal data of the target object; the second input channel is an input channel for abnormal features of the target object; and the third input channel is a global feature of visual image data of the target object.

8. The method according to claim 1, characterized in that The first detection model is trained based on target modal data of the target object and target modal data annotated with abnormal features of the target object.

9. The method according to claim 1, characterized in that: In the case where the target modality data is visual image data, the method further includes: Performing visualization processing on the abnormal information output by the first detection model to obtain a visualization image; The first detection model is optimized based on the difference between the visual image data of the target object and the visualized image.

10. The method according to claim 9, characterized in that The visualizing the abnormal information output by the first detection model to obtain a visual image includes: The abnormal information is visualized by using a gradient weighted class activation mapping method to obtain a visualized image.

11. The method according to claim 9 or 10, characterized in that: The visualization image is a thermal image.

12. The method according to claim 1, characterized in that The abnormal feature includes at least one of the following: abnormal category, abnormal region coordinates, and confidence score.

13. The method according to claim 1, characterized in that The scale of the second detection model is larger than that of the first detection model.

14. The method according to claim 1, characterized in that The target objects include automobile parts.

15. The method according to claim 1, characterized in that The target modality data includes at least one of the following: visual image data, sound data, and vibration data.

16. An anomaly detection method, characterized in that: The method comprises: Acquire target modal data of a target object to be detected; The target modal data is input into a second detection model to obtain abnormal information of the target object; wherein the second detection model is trained based on the model training method described in any one of claims 1 to 15.

17. The method according to claim 16, characterized in that The target objects include automobile parts.

18. The method according to claim 16, characterized in that The target modality data includes at least one of the following: visual image data, sound data, and vibration data.

19. An electronic device, characterized in that: It includes a processor and a memory, the processor is coupled to the memory; the memory is used to store computer instructions, and the computer instructions are loaded and executed by the processor to enable the computer device to implement the model training method as described in any one of claims 1 to 15.

20. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes computer execution instructions, and when the computer execution instructions are executed on a computer, the computer executes the model training method as described in any one of claims 1 to 15.

21. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is run on an electronic device, the electronic device executes the model training method as described in any one of claims 1 to 15.

Citation Information

Cited By

  • Cable production data management method and system based on Internet of Things

    CN120782337A