Data processing method and related device

WO2026200694A1PCT designated stage Publication Date: 2026-10-01HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/084714
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-26
Filing Date
2026-03-20
Publication Date
2026-10-01

Smart Images

  • Figure CN2026084714_01102026_PF_FP_ABST
    Figure CN2026084714_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Provided is a data processing method, comprising: a cloud management platform is capable of acquiring at least one reference image inputted by a user, the at least one reference image being used for describing a normal state of a target object; acquiring a first image to undergo inference inputted by the user, the first image being used for describing the target object; and invoking an AI model to perform inference on the at least one reference image and the first image, so as to obtain a first inference result regarding the first image, the first inference result being used for indicating whether the first image is in a normal state.
Need to check novelty before this filing date? Find Prior Art

Description

A data processing method and related equipment

[0001] This application claims priority to Chinese patent application filed on March 26, 2025, with application number CN202510371888.8 and title “A Data Processing Method and Related Equipment”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of artificial intelligence technology, specifically to a data processing method and related equipment. Background Technology

[0003] With the development of artificial intelligence (AI) technology, computer vision (CV) has achieved great success in industrial scenarios. It has been widely used in various scenarios such as industrial quality inspection, steel surface defect recognition, abnormal operation recognition, and personnel safety supervision, which has significantly reduced industrial costs and effectively improved industrial efficiency.

[0004] For example, in industrial applications, image processing models such as image detection models and image classification models are often used to perform inference tasks to identify defects in industrial products.

[0005] Taking industrial defect identification as an example, in actual training, the image processing model usually needs to be pre-trained in a supervised manner using labeled training data. Specifically, when collecting training data, the types of defects in industrial products need to be manually defined. Then, after collecting the training data, labels are assigned to the training data based on these defined defect types.

[0006] However, after supervised training, the types of defects that the trained image processing model can identify from industrial products are limited by the types of defects included in the labels of the training data. Since anomalies in industrial scenarios are often difficult to enumerate, the image processing model, constrained by the types of defects included in the labels of the training data, struggles to identify new categories such as novel anomalies. This limits the inference performance of the image processing model and affects detection accuracy. Summary of the Invention

[0007] This application provides a data processing method that can comprehensively detect anomalies and other situations in industrial scenarios using an AI model, achieving high detection accuracy. This application also provides corresponding devices, equipment, computer-readable storage media, and computer program products.

[0008] This application provides a data processing method applied to a cloud management platform. The cloud management platform manages infrastructure providing cloud services, including at least one cloud data center, and at least one server in the cloud data center is equipped with an AI model. The method includes: the cloud management platform acquiring at least one reference image input by a user, the reference image describing the normal state of a target object; the cloud management platform acquiring a first image to be inferred input by the user, the first image describing the target object; and the cloud management platform invoking an AI model to perform inference on the at least one reference image and the first image to be inferred, to obtain a first inference result regarding the first image to be inferred, the first inference result indicating whether the first image to be inferred is in a normal state, the first inference result being based on the difference between the first image to be inferred and the normal state of the target object.

[0009] In the first aspect, the AI ​​model can identify the features of the normal state of a target object using at least one reference image input by the user, and thus obtain a first inference result based on the difference between the first image to be inferred and the normal state of the target object. It is evident that the AI ​​model can identify the difference between the first image to be inferred and the normal state of the target object, thereby achieving the ability to identify "abnormal or abnormal," greatly improving the accuracy and comprehensiveness of identifying abnormal situations in practical application scenarios, and helping to promptly detect problems in various application scenarios such as industrial settings.

[0010] Furthermore, by identifying the difference between the first image to be inferred and the normal state of the target object, the condition of the first image to be inferred can be determined. This eliminates the need to predefine various types of abnormal situations in application scenarios such as industrial scenarios when labeling training data to obtain training data labels during the model training phase. This avoids the problem that the types of abnormal situations are difficult to fully define, which leads to the inability to fully identify abnormalities. Usually, it is only necessary to identify whether the training data is normal, which effectively reduces a lot of work in the labeling process of training data and improves the development efficiency of the training phase.

[0011] In one possible implementation of the first aspect, the first inference result is used to indicate that the first image to be inferred is not in a normal state, and the first inference result is also used to indicate the region in the first image to be inferred that is not in a normal state.

[0012] In one possible implementation, for example, if the target object is a part in an industrial scene, the reference image is used to describe a defect-free part, and the first image to be inferred includes images of parts produced on a production line, then the first inference result can indicate whether there is a defect on the part captured in the image of the first image to be inferred. Furthermore, if a defect exists, it indicates the location and / or outline of the defect, etc., to describe the area where the defect is located. It can be seen that the first inference result can be obtained based on the difference between the first image to be inferred and the normal state, thereby clearly indicating abnormal situations in the current scene to be inferred.

[0013] In one possible implementation of the first aspect, the AI ​​model includes a feature extraction layer and a feature processing layer. After acquiring at least one reference image input by the user, the model further includes: a cloud management platform extracting features from the at least one reference image through the feature extraction layer to obtain a first feature corresponding to the at least one reference image; the cloud management platform calling the AI ​​model to infer the at least one reference image and the first image to be inferred to obtain a first inference result about the first image to be inferred, including: the cloud management platform extracting features from the first image to be inferred through the feature extraction layer to obtain a second feature corresponding to the first image to be inferred; and the cloud management platform inferring the first feature and the second feature through the feature processing layer to obtain the first inference result.

[0014] In this possible implementation, after obtaining at least one reference image, regardless of whether subsequent images to be inferred (e.g., the first image to be inferred) are received, features can be extracted from at least one reference image in advance through the feature extraction layer to obtain the first features corresponding to at least one reference image, thereby shortening the inference time after obtaining the first image to be inferred.

[0015] In one possible implementation of the first aspect, the method further includes: acquiring a second image to be inferred input by the user; extracting features from the second image to be inferred through a feature processing layer to obtain a third feature corresponding to the second image to be inferred; and inferring from the first feature and the third feature through the feature processing layer to obtain a second inference result corresponding to the second image to be inferred.

[0016] In this possible implementation, the first feature can be used not only for the first inference task but also for the inference process corresponding to the second inference image in the second inference task. Therefore, the first feature corresponding to at least one reference image can be repeatedly used in the inference tasks of multiple sets of inference images for the user in the current application scenario. In practical applications, after obtaining the reference image, only one call to the feature extraction layer is needed to obtain the first feature, eliminating the need to repeatedly call the feature extraction layer to extract features from at least one reference image for each set of inference images. This significantly improves resource utilization and processing efficiency for each set of inference images.

[0017] In one possible implementation of the first aspect, after obtaining at least one reference image input by the user, the cloud management platform further includes: obtaining a third image to be inferred input by the user, the third image to be inferred being used to describe the target object; and, if the cloud management platform determines that at least one reference image is associated with the user, performing inference on at least one reference image and the third image to be inferred through an AI model to obtain a third inference result.

[0018] In this possible implementation, at least one reference image can be associated with the user, so that at least one reference image is bound to the user. Other users cannot use the reference image associated with the user to perform inference tasks without authorization, thus effectively ensuring the user's data security.

[0019] In one possible implementation of the first aspect, the AI ​​model includes a feature extraction layer and a feature processing layer. The feature processing layer includes one or more sub-processing layers, which are used to implement different types of inference tasks. The one or more types of inference tasks corresponding to the one or more sub-processing layers include one or more of the following: question answering tasks, classification tasks, segmentation tasks, detection tasks, and regression tasks. The method further includes: a cloud management platform acquiring instruction information, which is used to indicate a target type of inference task, and the target type of inference task belongs to one or more types of inference tasks; the cloud management platform calling the AI ​​model to perform inference on at least one reference image and a first image to be inferred, to obtain a first inference result about the first image to be inferred, including: the cloud management platform, according to the instruction information, performs inference on at least one reference image and the first image to be inferred through the feature extraction layer and the sub-processing layer corresponding to the target type of inference task in the AI ​​model, to obtain the first inference result.

[0020] In this possible implementation, the AI ​​model relies on its feature extraction capabilities from at least one reference image and a first image to be inferred. Through various sub-processing layers, the AI ​​model can perform one or more types of inference tasks. In other words, in this possible implementation, the AI ​​model can perform multiple types of inference tasks without having to build different AI models for each inference task independently. This reduces the difficulty of model construction, training, and deployment, and greatly improves the performance and resource utilization of the AI ​​model.

[0021] In practical applications, at least one sub-processing layer in the AI ​​model can be flexibly activated based on user-sent or pre-configured instructions to achieve the target type of inference task according to the actual scenario requirements, thus adapting to the needs of various application scenarios.

[0022] In one possible implementation of the first aspect, the AI ​​model is obtained by training based on one or more sets of training data pairs. Any set of training data pairs includes first training data, a first label corresponding to the first training data, second training data, and a second label corresponding to the second training data. The first label is used to indicate that a preset object in the first training data is in a normal state, and the second label is used to indicate that a preset object in the second training data is not in a normal state.

[0023] Unlike traditional supervised training, where training data labels need to include specific category information, this possible implementation uses a label for any training data point to indicate whether the data is in a normal state. In other words, when labeling training data, it is only necessary to distinguish whether the data corresponds to a normal state, without needing to predefine various different types of labels or spend a lot of effort distinguishing the types of training data. As can be seen, this possible implementation can eliminate a lot of labeling work, greatly saving related workload and improving the efficiency of model training development.

[0024] Furthermore, by training the AI ​​model to be trained using one or more sets of training data pairs, the AI ​​model in the iterative process can be instructed to learn to recognize the features of the same preset object in normal and abnormal states during the training process. This enables the trained AI model to distinguish the data corresponding to the normal state and the data corresponding to the abnormal state of the preset object. Thus, after training and deploying the AI ​​model, in actual applications, it can identify whether the target object in the image to be inferred is in a normal state based on a set of input data pairs describing the target object, such as the reference image and the image to be inferred.

[0025] A second aspect of this application provides a cloud management platform that has the functionality to implement the methods described in the first aspect or any possible implementation of the first aspect. This functionality can be implemented in hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the described functionality, such as an acquisition module and an inference module.

[0026] A third aspect of this application provides a computing device cluster including at least one computing device, the at least one computing device including a processor and a memory, the memory of the at least one computing device storing computer-executable instructions that can run on the processor, and when the computer-executable instructions are executed by the processor, the processor executes a method as described in the first aspect or any possible implementation of the first aspect.

[0027] The fourth aspect of this application provides a computer-readable storage medium storing one or more computer-executable instructions, wherein when the computer-executable instructions are executed by a processor, the processor performs a method as described in the first aspect or any possible implementation thereof.

[0028] The fifth aspect of this application provides a computer program product that stores one or more computer-executable instructions, wherein when the computer-executable instructions are executed by a processor, the processor executes a method as described in the first aspect or any possible implementation thereof.

[0029] A sixth aspect of this application provides a chip system including a processor for supporting the processor in implementing the functions involved in the first aspect or any possible implementation thereof. In one possible design, the chip system may further include a memory for storing necessary program instructions and data. This chip system may be composed of chips or may include chips and other discrete devices.

[0030] The technical effects of the second to sixth aspects or any of the possible implementations thereof can be found in the first aspect or the technical effects of the relevant possible implementations thereof, and will not be repeated here. Attached Figure Description

[0031] Figure 1 is an exemplary schematic diagram of a data center provided in an embodiment of this application;

[0032] Figure 2 is a schematic diagram of an exemplary system architecture provided in an embodiment of this application;

[0033] Figure 3 is a schematic diagram of an embodiment of the data processing method provided in this application;

[0034] Figure 4a is an exemplary flowchart of an AI model performing inference in a synchronous manner according to an embodiment of this application.

[0035] Figure 4b is an exemplary flowchart of an AI model performing inference asynchronously according to an embodiment of this application.

[0036] Figure 5 is an exemplary schematic diagram of the AI ​​model provided in an embodiment of this application;

[0037] Figure 6 is an exemplary schematic diagram of the reasoning process provided in an embodiment of this application;

[0038] Figure 7 is an exemplary schematic diagram of the AI ​​model provided in an embodiment of this application;

[0039] Figure 8 is an exemplary schematic diagram of the development process corresponding to the training process provided in the embodiments of this application;

[0040] Figure 9 is a schematic diagram of an embodiment of the cloud management platform provided in this application;

[0041] Figure 10 is a structural schematic diagram of a computing device provided in an embodiment of this application;

[0042] Figure 11 is a schematic diagram of a computing device cluster provided in an embodiment of this application;

[0043] Figure 12 is a schematic diagram of a computing device cluster provided in an embodiment of this application. Detailed Implementation

[0044] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.

[0045] As will be known to those skilled in the art, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0046] In this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms “comprising” and “having”, and any variations thereof, are intended to cover non-exclusive inclusion, such that a process, method, system, product, or apparatus that comprises a series of units is not necessarily limited to those units, but may include other units not expressly listed or inherent to those processes, methods, products, or apparatus.

[0047] 1. Large Language Model (LLM)

[0048] Large language models are deep learning models trained on massive amounts of text data that can generate natural language text or understand the meaning of language text. Large language models can handle various natural language tasks, such as text classification, question answering, and dialogue, and are an important pathway to artificial intelligence.

[0049] Specifically, large language models are a technology that has emerged in recent years. Because large language models undergo meticulous data engineering and training processes, their parameters have learned a wealth of existing natural language processing knowledge. This knowledge can now replace humans in many language-related tasks, such as having large language models write code or perform text summarization.

[0050] 2. Multimodal model

[0051] Multimodal models, such as the vision language model (VLM), aim to combine multimodal data, including visual and textual information, for reasoning. Currently, multimodal models have demonstrated enormous potential in various multimodal applications, such as visual reasoning, visual question answering, multimodal detection, and data classification.

[0052] With the development of AI technology, computer vision (CV) has achieved great success in industrial scenarios. It has been widely used in various scenarios such as industrial quality inspection, steel surface defect recognition, abnormal operation recognition, and personnel safety supervision, which has significantly reduced industrial costs and effectively improved industrial efficiency.

[0053] For example, in industrial applications, image processing models such as image detection models and image classification models are often used to perform inference tasks to identify defects in industrial products.

[0054] However, traditional solutions for inference tasks such as defect identification in industrial products using image processing models currently suffer from at least the following problems:

[0055] 1. Taking industrial defect identification as an example, in the actual training process, it is usually necessary to pre-train the image processing model in a supervised manner using corresponding labeled training data. Specifically, when collecting training data, the types of defects in industrial products need to be manually defined. Then, after collecting the training data, labels are assigned to the training data based on the defined defect types.

[0056] Thus, after supervised training, when performing defect identification tasks on industrial products using the trained image processing model, the types of defects identified by the trained image processing model from industrial products are limited to the types of defects contained in the labels of the training data.

[0057] However, anomalies in industrial scenarios are often difficult to enumerate. When image processing models are limited by the types of defects included in the training data labels, they cannot achieve the ability to identify "abnormal as abnormal," thus limiting their inference performance and affecting detection accuracy. Furthermore, in practical applications, the labeling of training data is tedious due to the large number of label types, resulting in high time and manpower costs.

[0058] It is evident that this traditional training method has a long training cycle, high cost, and poor scalability.

[0059] 2. When using image processing models for reasoning, the analysis is usually performed on a single image, which makes it difficult to make comprehensive inferences and limits the reasoning ability.

[0060] 3. In real-world reasoning scenarios, since there is no correlation between various reasoning tasks, it is necessary to construct corresponding image processing models for each reasoning task. For example, detection tasks require one image processing model, while classification tasks, segmentation tasks, etc., each require different image processing models, resulting in a large number of models and a large amount of resources required for construction and deployment.

[0061] Based on this, the embodiments of this application provide a data processing method that can detect anomalies and other situations in industrial scenarios in a relatively comprehensive manner through an AI model with high detection accuracy. In addition, the training workload of the AI ​​model is small, and in some examples, the AI ​​model can be easily applied to various inference tasks.

[0062] The method described in this application embodiment can be applied to a computing device cluster, which may include one or more computing devices.

[0063] The type of computing device is not limited here. For example, any computing device can be a terminal device, a server, a container, or a virtual machine, etc. Different computing devices can be of the same type or different types.

[0064] For example, in some examples, the cluster of computing devices can be terminal devices. Exemplarily, the terminal devices can be one or more of the following combinations: mobile phone, tablet, computer with wireless transceiver capabilities, virtual reality (VR) terminal, augmented reality (AR) terminal, terminal in industrial control, terminal in self-driving, terminal in remote medical care, terminal in smart grid, terminal in transportation safety, terminal in smart city, terminal in smart home, terminal in Internet of Things (IoT), wearable device, robot, etc.

[0065] In other examples, the cluster of computing devices can be used to implement a cloud management platform; in other words, the methods of this application embodiment can be applied to a cloud management platform.

[0066] A cloud management platform is used to manage the infrastructure that provides cloud services. It can provide computing, networking, and storage capabilities based on hardware and software resources. For example, the cloud management platform and infrastructure can reside in one or more data centers to provide cloud resources through those data centers.

[0067] The following is an exemplary description of a data center, illustrated in Figure 1.

[0068] In Figure 1, the cloud management platform interacts with one or more servers (Server 1 and Server 2 in Figure 1) through the data center's internal network. The servers consist of a hardware layer and a software layer. The hardware layer includes the server's hardware configuration, such as PCI devices like network interface cards (NICs), graphics processing units (GPUs), and offloading cards, which can be plugged into peripheral component interconnect (PCI) or peripheral component interconnect express (PCIe) slots. The software layer includes the operating system installed and running on the server (the operating system relative to the virtual machine can be called the host operating system). The host operating system contains a virtual machine manager (also called a hypervisor), whose role is to implement compute virtualization, network virtualization, and storage virtualization of the virtual machines and to manage them. A virtual machine (VM) refers to a complete computer system simulated by software, possessing full hardware system functionality and running in a completely isolated environment. In the system architecture shown in Figure 1, the infrastructure includes multiple servers, which can be used to run virtual machines. The specifications of the virtual machines can be the same or different. Virtual machines can also be called cloud servers (elastic compute service, ECS), elastic instances, etc., and different cloud service providers may have different names for them.

[0069] In one example of an embodiment of this application, the cloud management platform can be a public cloud platform. In this case, cloud service providers such as individuals or software developers with cloud resource development capabilities can provide cloud services to users. Users obtain cloud services through the Internet but do not own cloud computing resources. In other embodiments of this application, the cloud management platform can be a private cloud platform or a hybrid cloud platform, and this application does not impose any restrictions on this.

[0070] Specifically, in the example shown in Figure 1, the cloud management platform can provide an access interface (such as a user interface or application programming interface (API)). Users of the cloud management platform and cloud service providers can operate the client to remotely access the access interface to register a cloud account and password on the cloud management platform. After the cloud management platform successfully authenticates the cloud account and password, they can log in to the cloud management platform to create, manage, log in to and operate virtual machines in the cloud data center.

[0071] For example, referring to the example shown in Figure 2, users such as enterprises, organizations, or individuals can purchase cloud services to deploy AI models through cloud resources on a cloud management platform. When an AI model needs to perform inference tasks, the user can input reference information such as reference images and inference data such as the image to be inferred into the cloud management platform, and then invoke the AI ​​model to perform the corresponding inference task, obtaining the inference results from the cloud management platform. Furthermore, users can also use the cloud management platform to perform operations such as training the AI ​​model.

[0072] Of course, the cloud management platform can also be other types of cloud management platforms, and this application embodiment does not limit this.

[0073] Based on the cloud management platform described above, as shown in Figure 3, the data processing method of this application embodiment may include steps 301-303.

[0074] Step 301: The cloud management platform obtains at least one reference image input by the user.

[0075] At least one reference image is used to describe the normal state of the target object.

[0076] The specific target object can be determined based on the actual application scenario. The target object can be a physical object, such as a specified device or component; or it can be a digital object, such as software, a process, or a digital model like a 3D model or graph structure; or it can be an abstract object, such as a work process. Therefore, the target object can have various forms, and no limitation is made here.

[0077] The at least one reference image can be used to describe the normal state of the target object, such as normal shape, normal working state, normal process, etc.

[0078] The specific content and form of this normal state can be determined based on the type of the target object and the actual application scenario. The at least one reference image describing this normal state can have various forms, which are not limited here. For example, when multiple reference images exist, these multiple reference images can be in the form of a video, specifically multiple video frames; or, they can be an image sequence; or, each of the multiple reference images can be a single image. The data type of any reference image is not limited here, and when there are multiple reference images, the data types of different reference images can be the same or different.

[0079] In addition, in some examples, the user can input other reference information related to the at least one reference image. In other words, in some examples, the user can input reference information, which may include, but is not limited to, the at least one reference image. For example, the reference information may also include one or more other reference information such as text data, voice data, and graph data, so as to describe the normal state of the target object through one or more modal reference information.

[0080] For example, if the target object is a physical object, such as equipment, a part, or a product produced on an assembly line, then the normal state of the target object can indicate one or more of the following: normal shape, normal motion state, etc. In this case, the reference information can describe the normal shape and / or normal motion state of the target object through data from one or more modalities.

[0081] For example, the target object is a part produced in an industrial manufacturing process. The normal state of this part indicates that it is defect-free, and at least one reference image of the part can include a series of images of the defect-free part. For instance, it could include two-dimensional images taken from different angles of the defect-free part, or three-dimensional images obtained by scanning the defect-free part. Furthermore, the user can input descriptive text about the defect-free part, as well as textual descriptions of the series of images of the defect-free part, and other reference information related to the at least one reference image.

[0082] If the target object is a digitized object, such as software or a digitized model, then the normal state of the target object can indicate that the digitized content or specified indicators involved in the digitized object conform to the specified specifications. For example, if the target object is a digitized 3D model, then the normal state of the target object can indicate that the digitized form of the target object is in a normal form, and at least one reference image of the digitized 3D model can describe the various 3D sub-models contained in the digitized model. In addition, the user can also input other reference information related to the at least one reference image, such as the 3D sub-model itself, textual or graphical descriptions of the connection relationships between the various 3D sub-models, and / or textual descriptions of the constraints corresponding to the digitized model, etc.

[0083] If the target object is an abstract object, such as a work process, then the normal state of the target object can indicate the normal work process of the target object. For example, if the target object is a work process for a specific business, then the normal state of the work process for the specified business can indicate that the work process for the specified business conforms to the normal work process, and at least one reference image for the work process of the specified business can describe each work node of the normal work process. In addition, the user can also input other reference information related to the at least one reference image, such as a graph structure describing the relationship between work nodes, and may also include textual descriptions of the graph structure and / or the normal work process.

[0084] As can be seen, the target object and its normal state can be flexibly determined according to the actual application scenario. The data form and specific content of the reference information, such as at least one reference image describing the normal state of the target object, can also be varied and are not limited here.

[0085] The reference information, such as at least one reference image, can be input by the user through an application programming interface (API), a webpage, or other means. Alternatively, the user can select at least one reference image from a set of candidate reference images pre-input by the user through a configuration interface. In this embodiment, the specific method by which the user inputs the reference information is not limited.

[0086] It should be noted that in the embodiments of this application, the user can be a registered user or an unregistered user. For example, in some examples, the method of the embodiments of this application is implemented through designated software. In this case, the user can be a registered and logged-in user of the designated software, thereby transmitting the at least one reference image to the software. Alternatively, the user can also be an unregistered user trying out the functions of the designated software as a guest, and thus can also input the at least one reference image into the designated software.

[0087] Step 302: The cloud management platform obtains the first image to be inferred input by the user.

[0088] The first image to be inferred is used to describe the target object.

[0089] In this embodiment, the first image to be inferred is the data to be inferred corresponding to the task to be processed regarding the target object in a real-world application scenario. This first image to be inferred may include one or more images.

[0090] The task to be processed can also be considered a reasoning task, and the specific type of the reasoning task can be various. For example, the reasoning task can include one or more of the following types of tasks: classification task, detection task, segmentation task, and regression task.

[0091] As can be seen, in this embodiment of the application, the inference task may include one task or multiple tasks. When the inference task includes one or multiple tasks, the execution method of each task based on the AI ​​model can be referred to step 303 and the subsequent related embodiments, and will not be repeated here.

[0092] The first image to be inferred can be considered as being contained within the first data to be inferred. It is understood that the first data to be inferred input by the user includes, but is not limited to, the first image to be inferred. The data types included in the first data to be inferred can also be various; for example, the first data to be inferred may also include descriptive text corresponding to the first image to be inferred.

[0093] The data type of the first data to be inferred can be the same as or different from the data type of the reference information. For example, the data type of the first data to be inferred can be all or part of the data types in the reference information. In addition, the data type of the first data to be inferred can also include data types not mentioned in the reference information. For example, the reference information includes at least one reference image and descriptive text describing the at least one reference image, while the first data to be inferred includes the first image to be inferred but does not include the descriptive text of the first image to be inferred.

[0094] There are various ways to obtain the first image to be inferred, and no limitation is made here. For example, the first image to be inferred may be locally stored in the computing device cluster executing the embodiments of this application, or it may be transmitted to the computing device cluster by other devices such as client devices, or it may be collected by the computing device cluster through a webpage or other applications, or it may be obtained by the computing device cluster after inferring specified text data, image data, voice data, etc.

[0095] It should be noted that, in the embodiments of this application, the order in which at least one reference image and the first image to be inferred are acquired is not restricted.

[0096] In some examples, a computing device cluster executing the methods of this application embodiment can obtain at least one reference image and a first image to be inferred from the same data packet received by the same API. Alternatively, it can obtain at least one reference image and the first image to be inferred sequentially through the same API; for example, it can obtain at least one reference image first and then obtain the first image to be inferred through the same API. Or, the computing device cluster executing the methods of this application embodiment can obtain at least one reference image and the first image to be inferred separately through different methods. For example, a user can send at least one reference image from a client device to the computing device cluster of this application embodiment through a first API, and then send the first image to be inferred from the client device to the computing device cluster of this application embodiment through a second API.

[0097] As can be seen, there are various ways and orders in which the computing device cluster executing the method of this application embodiment acquires the reference image and the first image to be inferred. This application embodiment does not limit the execution order between steps 301 and 302. In other words, steps 301 and 302 can be executed sequentially, or they can be implemented through the same acquisition operation, or steps 302 can be executed first, and then steps 301 can be executed.

[0098] Step 303: The cloud management platform calls the AI ​​model to perform reasoning on at least one reference image and the first image to be reasoned, so as to obtain the first reasoning result about the first image to be reasoned.

[0099] The first inference result is used to indicate whether the first image to be inferred is in a normal state. The first inference result is obtained based on the difference between the first image to be inferred and the normal state of the target object.

[0100] In this embodiment of the application, the specific type of the AI ​​model can vary depending on the actual application scenario requirements.

[0101] In some examples, the data types involved, such as the reference information and the first data to be inferred, are single-modal data. In such cases, the AI ​​model is used to perform inference on that single-modal data. For example, if the reference information includes only at least one reference image and the first data to be inferred includes only the first image to be inferred, then the AI ​​model can be an image processing model.

[0102] In other examples, the data type involved is multimodal data, meaning that the reference information and the first data to be inferred involve multiple modalities. In such cases, the AI ​​model can be a multimodal model. For instance, when the reference information and the first data to be inferred include images and text, the AI ​​model can be an existing or future-developed visual-language model (VLM). Of course, when the reference information and the first data to be processed involve data including other modalities, the AI ​​model can also be a multimodal model corresponding to other types of multimodal data.

[0103] The AI ​​model possesses the ability to understand the features of reference information and the first data to be inferred. Specifically, it has the ability to understand the features of at least one reference image and the first image to be inferred. This means that the AI ​​model can identify and extract the feature information of the reference image and the first image to be inferred, and then perform inference based on the feature information of the reference image and the first image to be inferred, such as classification, anomaly detection, and segmentation.

[0104] The AI ​​model's ability to understand the features of reference information and the first inference data, such as at least one reference image and the first image to be inferred, can be achieved by training with a large amount of training data. For a related introduction to the training of the AI ​​model, please refer to the subsequent related embodiments, which will not be repeated here.

[0105] After obtaining at least one reference image and a first image to be inferred, an AI model can be used to infer from the at least one reference image and the first image to be inferred to obtain a first inference result about the first image to be inferred.

[0106] At least one reference image can be directly input into the AI ​​model for inference, or it can undergo specified preprocessing operations before being input into the AI ​​model. For example, at least one reference image can be preprocessed by performing format conversion, data cleaning, or other preprocessing operations before being input into the AI ​​model. The first image to be inferred can be directly input into the AI ​​model for inference, or it can undergo specified preprocessing operations before being input into the AI ​​model. For example, the first image to be inferred can be preprocessed by performing format conversion, data cleaning, or other preprocessing operations before being input into the AI ​​model.

[0107] The AI ​​model can process at least one reference image and the first image to be inferred in multiple ways. For example, the AI ​​model can perform inference on at least one reference image and the first image to be inferred synchronously or asynchronously. The specific processes of AI model inference in synchronous and asynchronous ways are described below.

[0108] 1. The AI ​​model uses a synchronous approach for inference.

[0109] This example illustrates the specific implementation of synchronous reasoning in an AI model, using the reasoning process of at least one reference image and a first image to be reasoned as an example. Specifically, Figure 4a shows an exemplary flowchart of an AI model performing synchronous reasoning on at least one reference image and a first image to be reasoned.

[0110] Referring to the example shown in Figure 4a, in some examples, during the data transmission phase, the user can transmit reference information about the target object and the first inference data from the client device to the computing device cluster via a specified API.

[0111] For example, in the example shown in Figure 4a, a user can input reference information including at least one reference image and related textual descriptions of the at least one reference image, and can also input first inference data including a first image to be inferred. Of course, in other examples, the user can also use other methods to send the reference information and first inference data, such as at least one reference image and a first image to be inferred, to the computing device cluster. For example, the user can send the reference information and first inference data to the computing device cluster through a visual interface.

[0112] During the inference phase, after receiving information such as at least one reference image and a first image to be inferred, the computing device cluster can input this information as input data into the AI ​​model. The AI ​​model then performs an inference task on this information, outputting a first inference result regarding the first image to be inferred. For example, a feature extraction layer in the AI ​​model can extract first features from the at least one reference image and second features from the first image to be inferred. A feature processing layer in the AI ​​model can then perform inference on the first and second features to output the first inference result.

[0113] In this way, during the result return phase, the computing device cluster can send the first inference result about the first image to be inferred to the user's client device through the API. For example, the first inference result can indicate whether the first image to be inferred contains information such as defects or anomalies. Of course, the first inference result can also include other information. For details, please refer to the subsequent introduction about the first inference result.

[0114] 2. The AI ​​model uses an asynchronous approach for inference.

[0115] This example illustrates the implementation of a synchronous reasoning method for an AI model, using the reasoning process involving at least one reference image and a third image to be reasoned. Specifically, in some embodiments, it further includes:

[0116] The cloud management platform obtains user information entered by the user;

[0117] Following step 301, the following is also included:

[0118] The cloud management platform associates at least one reference image with the user;

[0119] The method includes:

[0120] The cloud management platform obtains the third image to be inferred from the user's input. The third image to be inferred is used to describe the target object.

[0121] Once the cloud management platform determines that at least one reference image is associated with the user, it uses an AI model to infer from at least one reference image and a third image to be inferred, in order to obtain a third inference result.

[0122] In this embodiment, the third image to be inferred can refer to the relevant description of the first image to be inferred, and the third inference result can refer to the relevant description of the first inference result, which will not be repeated here. In actual application, the third image to be inferred can be a different image from the first image to be inferred; or, the first image to be inferred can be used as the third image to be inferred. In this case, it can be considered that the AI ​​model executes step 303 asynchronously.

[0123] Below, for ease of description, we can take the third image to be inferred as the first image to be inferred as an example to introduce the specific process of the AI ​​model executing step 303 in an asynchronous manner, so as to introduce the specific way of the AI ​​model to perform inference in an asynchronous manner. When the third image to be inferred is a different image from the first image to be inferred, the specific way of obtaining the third inference result by the AI ​​model inferring from at least one reference image and the third image to be inferred can be referred to the specific process of the AI ​​model executing step 303 in an asynchronous manner in this example.

[0124] In this embodiment of the application, the user information can be used to identify the user. For example, it may include one or more of the user's name, mobile phone number, email address, etc. In addition, the user information may also include other information of the user, such as one or more of the user's address, preferences, etc., without limitation.

[0125] In this embodiment of the application, the user can transmit at least one reference image and a first image to be inferred through two data transmissions.

[0126] Specifically, Figure 4b shows an exemplary flowchart of an AI model acquiring and inferring at least one reference image and a first image to be inferred in an asynchronous manner.

[0127] Referring to the example shown in Figure 4b, when a user registers with a cloud management platform implemented by a computing device cluster via an API, during the information registration phase, the user transmits at least one reference image to the computing device cluster via the API to indicate that the at least one reference image is associated with the user. This allows the computing device cluster to save the user's information and the corresponding at least one reference image during the information registration phase. At this time, the association between the identifier of the at least one reference image (e.g., the name, image storage path, storage address, etc.) and the user's identifier (e.g., the user's name, ID, etc.) can be stored in a database or a specified list. Alternatively, a tag can be added to the at least one reference image, which may include the user's identifier to indicate that the at least one reference image is associated with the user. Therefore, there are multiple ways to associate at least one reference image with a user in the cloud management platform, and no limitation is made here.

[0128] Furthermore, in other examples, after a user registers by entering their user information (i.e., after the information registration phase) and before transmitting the first image to be inferred, the computing device cluster can then transmit at least one reference image to the computing device cluster. The computing device cluster can then associate this at least one reference image with the user, binding it to that user and preventing other users from using the at least one reference image associated with that user without authorization, effectively ensuring user data security.

[0129] In this scenario, before executing an actual inference task, a user can pre-transmit and store at least one reference image describing the target object in the current scene to the computing device cluster. After this pre-transmission, the user does not need to repeatedly transmit the at least one reference image to the computing device cluster for each inference task. Instead, the user can instruct the computing device cluster to reuse the pre-transmitted and stored reference image in subsequent inference tasks for the target object, effectively improving data utilization, reducing resource consumption, and increasing processing efficiency.

[0130] For example, in some cases, after obtaining at least one reference image and the first image to be inferred, the at least one reference image and the first image to be inferred can be input together into the AI ​​model for inference to obtain the first inference result.

[0131] In other examples, during the time period after at least one reference image is transmitted to the computing device cluster but before the first image to be inferred is transmitted to the computing device cluster, the AI ​​model can use a structure that can independently process at least one reference image to infer at least one reference image in advance and obtain the first feature corresponding to at least one reference image.

[0132] Specifically, referring to the example shown in Figure 5, in some embodiments, the AI ​​model includes, but is not limited to, a feature extraction layer and a feature processing layer.

[0133] The feature extraction layer is used to extract features from the input data. In many scenarios, when input data from different modalities are used for inference through the feature extraction layer, feature fusion between multimodal data is not required. Therefore, reference information including at least one reference image and first inference data including the first image to be inferred can be processed independently by the feature extraction layer. Furthermore, it should be noted that the function of the feature extraction layer is not limited to extracting features from input data such as at least one reference image and the first image to be inferred; for example, the feature extraction layer can also be used to encode the input data. The feature extraction layer may include one or more layers, and the specific structure of the feature extraction layer is not limited here.

[0134] Considering the different modalities of the reference information and the first data to be inferred, the feature extraction layer can have various structures, which are not limited here.

[0135] The following section uses an AI model as an example to illustrate the structure of the feature extraction layer.

[0136] It should be noted that the number of sub-feature extraction layers (such as the subsequent first and second feature extraction layers) and their corresponding modalities in this feature extraction layer can vary. This example is for illustrative purposes only and is not a limitation.

[0137] In one example, referring to the example shown in Figure 5, the feature extraction layer includes a first feature extraction layer and a second feature extraction layer. The first feature extraction layer is used to extract image features and can also be considered as an image encoder. The second feature extraction layer is used to extract text features and can also be considered as a text encoder.

[0138] The feature extraction layer can independently reason about the reference information without relying on the first data to be reasoned, so as to extract the first feature of the reference information. Since the reference information includes at least one reference image, the first feature can also be considered as the first feature corresponding to at least one reference image.

[0139] For example, when the reference information includes at least one reference image and corresponding text description information, inference can be performed on at least one reference image through a first feature extraction layer, and inference can be performed on the text description information corresponding to at least one reference image through a second feature extraction layer. The outputs of the first and second feature extraction layers are then fused to obtain the first feature of the reference information. The features obtained by inferring from at least one reference image through the first feature extraction layer and the features obtained by inferring from the text description information corresponding to at least one reference image through the second feature extraction layer can be features mapped to the same feature space, thus facilitating feature fusion and other processing in subsequent feature processing layers.

[0140] Furthermore, the feature extraction layer can also independently infer the first data to be inferred in order to extract the second feature of the first data to be inferred. Since the first data to be inferred includes the first image to be inferred, the second feature can also be considered as the second feature corresponding to the first image to be inferred.

[0141] The feature processing layer is used to fuse the first feature and the second feature to obtain the first inference result. The specific structure of this feature processing layer is not limited here. For example, the feature processing layer may include one or more of the following structures: expert module, self-attention mechanism, feedforward neural network (FFN) residual connection, or layer normalization.

[0142] Furthermore, in some examples, before the computing device cluster obtains the first inference data such as the first image to be inferred, the reference information such as at least one reference image can be inferred in advance through the feature extraction layer, thereby shortening the data processing time after obtaining the first inference data such as the first image to be inferred.

[0143] Specifically, in some embodiments, after acquiring at least one reference image input by the user, the method further includes:

[0144] The cloud management platform extracts features from at least one reference image through a feature extraction layer to obtain the first feature corresponding to at least one reference image;

[0145] The cloud management platform invokes an AI model to perform reasoning on at least one reference image and a first image to be inferred, in order to obtain a first inference result regarding the first image to be inferred, including:

[0146] The cloud management platform extracts features from the first image to be inferred through a feature extraction layer to obtain the second feature corresponding to the first image to be inferred.

[0147] The cloud management platform uses a feature processing layer to reason about the first and second features to obtain the first reasoning result.

[0148] Referring to the example shown in Figure 6, in this embodiment of the application, after obtaining at least one reference image, regardless of whether subsequent images to be inferred (such as the first image to be inferred or the second image to be inferred in subsequent examples) are received, features can be extracted from at least one reference image in advance through the feature extraction layer to obtain the first feature corresponding to at least one reference image.

[0149] Then, in subsequent applications, if the first image to be inferred is obtained, the first feature can be directly obtained, and the feature extraction layer can be called to extract features from the first image to be inferred. Then, the first feature and the second feature can be inferred through the feature processing layer to obtain the first inference result.

[0150] In one example, referring to the example shown in Figure 4b, after the user acquires a reference image and obtains the first feature by performing feature extraction on the reference image through the feature extraction layer, the user can save the first feature corresponding to the reference image.

[0151] In the subsequent inference task, during the data transmission phase, after the user transmits the first image to be inferred to the computing device cluster, if the computing device cluster determines that the first image to be inferred is associated with the user, it retrieves the pre-saved first feature, extracts features from the first image to be inferred through the feature extraction layer, obtains the corresponding second feature, and then performs inference on the first and second features through the feature processing layer to obtain the first inference result. This first inference result can then be returned to the user.

[0152] Alternatively, in the inference task, if the computing device cluster determines that the first image to be inferred is not associated with any user, such as being transmitted by an unregistered customer, then the computing device cluster cannot obtain the first feature and cannot use the AI ​​model in this embodiment for inference, effectively ensuring the data security of registered users. In this case, referring to the example shown in Figure 4b, the computing device cluster can provide the user with other algorithms (such as traditional image processing models or other preset traditional algorithms) for the user to infer the first image to be inferred and obtain the corresponding inference results.

[0153] Furthermore, in some examples, the first feature corresponding to the at least one reference image stored in the computing device cluster can be repeatedly used in the current application scenario for multiple sets of inference data of the user.

[0154] Specifically, in some embodiments, the method further includes:

[0155] The cloud management platform acquires the second image to be inferred, input by the user;

[0156] The cloud management platform extracts features from the second image to be inferred through a feature processing layer to obtain the third feature corresponding to the second image to be inferred.

[0157] The cloud management platform uses a feature processing layer to reason about the first and third features to obtain the second reasoning result corresponding to the second image to be reasoned.

[0158] The second image to be reasoned is similar to the first image to be reasoned. For details, please refer to the relevant examples of the first case to be processed, which will not be repeated here.

[0159] Referring to the example shown in Figure 6, features can be extracted from the second image to be inferred through a feature extraction layer to obtain a third feature corresponding to the second image to be inferred. For example, a user can input second inference data including text description information corresponding to the second image to be inferred. Then, the first feature extraction layer extracts features from the second image to be inferred, and the second feature extraction layer extracts features from the text description information corresponding to the second image to be inferred. The outputs of the first and second feature extraction layers are then fused to obtain the third feature. Then, the feature processing layer can perform inference on the first and third features to obtain the second inference result corresponding to the second image to be inferred.

[0160] As can be seen, in this embodiment, the first feature can be used not only in the first inference task, specifically in the processing of the first image to be inferred in the first inference task, but also in the processing of the second image to be inferred in the second inference task. Therefore, the first feature corresponding to the reference image can be repeatedly used in the inference tasks of multiple sets of images to be inferred by the user in the current application scenario. In practical application scenarios, after obtaining the reference image, only one call to the feature extraction layer is needed to obtain the first feature, without repeatedly calling the feature extraction layer to extract features from the reference image for each set of images to be inferred, greatly improving resource utilization and processing efficiency for each set of images to be inferred.

[0161] Through any of the above embodiments, an AI model can be used to infer at least one reference image and a first image to be inferred, in order to obtain a first inference result about the first image to be inferred. The specific content of the first inference result can be varied and is not limited here.

[0162] For example, the first inference result is used to indicate whether the first image to be inferred is in a normal state, and when the first image to be inferred is not in a normal state, the first inference result is used to indicate the area in the first image to be inferred that is not in a normal state.

[0163] For example, if the target object is a part in an industrial setting, at least one reference image is used to describe a defect-free part, and the first inference image includes images of parts produced on a production line, then the first inference result can indicate whether there is a defect on the part captured in the first inference image. Furthermore, if a defect exists, it can also indicate the location and / or outline of the defect. Thus, the first inference result can be obtained based on the difference between the first inference image and the normal state of the target object. Specifically, based on the description of step 303 above, the AI ​​model can extract features from the first inference image and features from at least one reference image, thereby determining whether the first inference image conforms to the normal state of the target object based on the difference between the features of the first inference image and the features of at least one reference image, to obtain the first inference result. Thus, the first inference result is obtained based on the difference between the features of the first inference image and the features of at least one reference image, and can reflect whether the first inference image conforms to the normal state of the target object. The first inference result can directly indicate the parts in the first inference image that differ from the normal state of the target object, as areas in the first inference image that are not in a normal state. Alternatively, the difference between the first image to be inferred and the normal state of the target object can be further processed to obtain the first inference result; for example, in some examples, the areas in the first image to be inferred that are not in a normal state (that is, in an abnormal state) can be subjected to other enhancement operations such as magnification to obtain the first inference result.

[0164] In this embodiment, the first reasoning result is obtained based on the difference between the first image to be reasoned and the normal state of the target object. It can be seen that the AI ​​model can identify the difference between the first image to be reasoned and the normal state of the target object, thereby achieving the ability to identify abnormalities as "abnormal or not". This greatly improves the accuracy and comprehensiveness of identifying abnormal situations in practical application scenarios and helps to discover problems in a timely manner in various application scenarios such as industrial scenarios.

[0165] Furthermore, by identifying the difference between the first image to be inferred and the normal state of the target object, the condition of the first image to be inferred can be determined. This eliminates the need to predefine abnormal situations in application scenarios such as industrial scenes when labeling training data to obtain training data labels during the model training phase. This avoids the problem that the types of abnormal situations are difficult to fully define, which leads to the inability to fully identify abnormalities. Usually, it is only necessary to identify whether the training data is normal, which effectively reduces a lot of work in the labeling process of training data and improves the development efficiency of the training phase.

[0166] Furthermore, in this embodiment, the specific content of the first reasoning result can be determined based on the reasoning task corresponding to the first image to be reasoned.

[0167] Specifically, in some embodiments, the AI ​​model includes a feature extraction layer and a feature processing layer. The feature processing layer includes one or more sub-processing layers. Different sub-processing layers are used to implement different types of inference tasks. The one or more types of inference tasks corresponding to one or more sub-processing layers include one or more of the following: question answering tasks, classification tasks, segmentation tasks, detection tasks, and regression tasks.

[0168] The method also includes:

[0169] The cloud management platform obtains instruction information, which is used to indicate the inference task of the target type. The inference task of the target type belongs to one or more types of inference tasks.

[0170] The cloud management platform invokes an AI model to perform reasoning on at least one reference image and a first image to be inferred, in order to obtain a first inference result regarding the first image to be inferred, including:

[0171] Based on the instructions, the cloud management platform uses the feature extraction layer in the AI ​​model and the sub-processing layer corresponding to the inference task of the target type to perform inference on the reference image and the first image to be inferred, so as to obtain the first inference result.

[0172] In this embodiment of the application, considering the correlation between different inference tasks for the same data, for example, in an industrial quality inspection scenario, in addition to detecting the defects of the corresponding object, it is also desirable to determine the location and area of ​​the defects through segmentation and other methods.

[0173] Based on this, in the embodiments of this application, the output of the feature processing layer in the AI ​​model may include one or more branches, which may perform their respective inference tasks based on the output of the feature extraction layer of the AI ​​model preceding the branch and other layers in the feature processing layer such as the self-attention layer.

[0174] Specifically, the feature processing layer may include, but is not limited to, one or more sub-processing layers. Different sub-processing layers are used to implement different types of reasoning tasks. The one or more types of reasoning tasks corresponding to one or more sub-processing layers include one or more of the following: classification tasks, segmentation tasks, detection tasks, and regression tasks.

[0175] The number of sub-processing layers is not limited here; they can be pre-constructed and trained according to user needs, etc.

[0176] For example, referring to the example shown in Figure 7, in one example, the feature processing layer may include three sub-processing layers: sub-processing layer 1 for performing a question-answering task, sub-processing layer 2 for performing a defect detection task, and sub-processing layer 3 for performing a defect segmentation task. Sub-processing layer 1 can output the answer to the question, sub-processing layer 2 can output defect location information such as the location coordinates of the defect, and sub-processing layer 3 can output the outline and / or area of ​​the defect.

[0177] Furthermore, before the three sub-processing layers of this feature processing layer, other layers such as a self-attention layer may be included, for example, to fuse multimodal features for inference. The structure of any sub-processing layer is not limited here. For example, any sub-processing layer may include a fully connected layer or similar structures. It is understood that in this embodiment, inferring input data such as at least one reference image, the textual description information of at least one reference image, and the first image to be inferred through the feature extraction layer and the sub-processing layer corresponding to the inference task of the target type in the AI ​​model is not limited to using only the feature extraction layer and the sub-processing layer corresponding to the inference task of the target type for inference; other layers in the AI ​​model may also be used for inference, such as the self-attention layer in the feature processing layer.

[0178] In this way, relying on the AI ​​model's ability to understand the features of at least one reference image and the first image to be inferred, the AI ​​model can perform one or more types of inference tasks through each sub-processing layer. That is to say, the AI ​​model of this application embodiment can realize multiple types of inference tasks without having to build different AI models independently for each inference task, reducing the difficulty of model construction, training and deployment, and greatly improving the performance and resource utilization of the AI ​​model.

[0179] In practical applications, at least one sub-processing layer in the AI ​​model can be flexibly activated based on the instruction information to achieve the target type of reasoning task according to the actual scenario requirements, obtain the corresponding reasoning results, and adapt to the needs of various application scenarios.

[0180] The indication information can be sent by the user to the computing device cluster, or it can be pre-configured in the computing device cluster, or it can be generated based on factors such as the resource status of the computing device cluster (e.g., whether it is currently idle).

[0181] For example, in the example shown in Figure 7, the user can select a defect segmentation task through a visual interface, and then the computing device cluster can receive the instruction information indicating the defect segmentation task.

[0182] Furthermore, users can send reference information, including at least one reference image and corresponding text description information, as well as a first image to be inferred to the computing device cluster to trigger defect detection and defect segmentation tasks.

[0183] When the reference information includes at least one reference image and corresponding text description information, inference can be performed on at least one reference image through a first feature extraction layer, and inference can be performed on the text description information corresponding to at least one reference image through a second feature extraction layer. The outputs of the first and second feature extraction layers are used as the first features corresponding to at least one reference image. Furthermore, features can be extracted from the first image to be inferred through the first feature extraction layer to obtain the second features corresponding to the first image to be inferred. By fusing the first and second features, the output of the feature extraction layer is obtained, and this output is input into the feature processing layer.

[0184] Then, through the processing of one or more intermediate layers such as the self-attention layer in the feature processing layer, the sub-processing layer 2 and sub-processing layer 3 are called for inference to obtain the output of sub-processing layer 2 and the output of sub-processing layer 3 as the first inference result. The first inference result can indicate whether there is a defect in the part described by the first image to be inferred, and when there is a defect, the area of ​​the defect, the position coordinates of the defect, and the segmented contour are output.

[0185] The training process of the AI ​​model in this application embodiment is described below.

[0186] The AI ​​model in this application embodiment can be trained by a computing device cluster executing any of the above embodiments, or it can be trained by other devices and then called by the computing device cluster executing any of the above embodiments. This application embodiment does not limit this.

[0187] For example, AI models can be trained in a supervised or semi-supervised manner, and the specific training process of the AI ​​model is not limited here.

[0188] In some embodiments, in order to ensure better training results, one or more sets of training data pairs containing labels can be used to train the AI ​​model in a supervised or semi-supervised manner.

[0189] Specifically, in some embodiments, the AI ​​model is obtained by training based on one or more sets of training data pairs. Any set of training data pairs includes first training data, a first label corresponding to the first training data, second training data, and a second label corresponding to the second training data. The first label is used to indicate that a preset object in the first training data is in a normal state, and the second label is used to indicate that a preset object in the second training data is not in a normal state.

[0190] The training data set can be a partial training data set or the complete training data set for the AI ​​model.

[0191] Furthermore, unlike traditional supervised training where training data labels need to include specific category information, in this embodiment, the label corresponding to any training data is used to indicate whether the corresponding training data is in a normal state. In other words, when labeling training data, it is only necessary to distinguish whether the training data corresponds to a normal state, without needing to predefine various different types of labels, nor is it necessary to spend a lot of effort distinguishing the types of training data during the labeling process. It can be seen that this training data labeling method in this embodiment can eliminate a lot of labeling work, greatly save related workload, and improve the efficiency of model training-related development work.

[0192] In this way, by training the AI ​​model with one or more sets of training data, the AI ​​model can be instructed to learn to recognize the features of the same preset object in normal and abnormal states during the training process. This enables the trained AI model to distinguish the data corresponding to the normal state and the data corresponding to the abnormal state of the preset object. Thus, after training and deploying the AI ​​model, in actual applications, it can identify whether the target object in the image to be inferred is in a normal state based on a set of input data pairs describing the target object, such as the reference image and the image to be inferred.

[0193] The following section uses images as an example to introduce the development process corresponding to the traditional training process and the development process corresponding to the training process in the embodiments of this application.

[0194] Referring to the example shown in Figure 8, when the training data includes images, the traditional training process involves installing a camera at a designated location, capturing images, and then copying the captured images to the development environment. After the images are copied to the development environment, they need to be analyzed and categorized to determine the various possible label contents.

[0195] For example, when dealing with defects in industrial products, the traditional training process requires manual traversal of all images to identify various defect categories in order to cover as many defects as possible. Therefore, this process usually takes a certain amount of time. For instance, in some examples, when the number of images is around 5,000, it may take about 3 days.

[0196] Then, in the traditional training process and its corresponding development workflow, the images used as training data can be labeled according to the various categories obtained from the analysis and organization. Since the number of images used as training data is usually large, the time required for this data labeling is usually long. For example, in some cases, when the number of images is around 5,000, it takes about 15 days.

[0197] In addition to data labeling, long-tail data collection is required to supplement the training data, ensuring that it covers as many defect categories of industrial products as possible and that there is a sufficient amount of training data for each defect category. In some examples, when the number of images is around 5,000, this supplementary collection process can take about 10 days.

[0198] Then, the corresponding labeled training data can be obtained for subsequent steps such as model training.

[0199] As can be seen, in some examples, when the number of images is around 5,000, it takes about 20 days from analyzing and organizing the categories of labels to obtaining the corresponding labeled training data. This incurs significant manpower and time costs, and the labeling process is also quite cumbersome.

[0200] In this solution, after the captured images are copied to the development environment, it is only necessary to distinguish whether the images correspond to a normal state to label the images and obtain corresponding training data pairs. Each training data pair can include an image describing the normal state of a preset object and an image describing the abnormal state of that preset object. In some examples, when the number of images is around 5000, the data differentiation and labeling process takes approximately 1.5 days.

[0201] As can be seen, in this embodiment of the application, the data differentiation and labeling process saves a lot of work compared to the traditional training process, greatly reducing the development workload and significantly improving the efficiency of the development process.

[0202] Furthermore, in some examples, labeled training data can be obtained for training sub-processing layers of the AI ​​model, and the performance of these sub-processing layers in performing corresponding inference tasks can be ensured through supervised training or fine-tuning. For example, for a sub-processing layer performing a question-answering task, the labels of its training data include preset answers; for a sub-processing layer performing a defect detection task, the labels of its training data include preset defect location coordinates and other location information; and for a sub-processing layer performing a defect segmentation task, the labels of its training data include preset defect contour information.

[0203] The data processing method provided by the embodiments of this application has been described above from multiple aspects. The cloud management platform provided by the embodiments of this application will be described below with reference to the accompanying drawings.

[0204] As shown in Figure 9, this application embodiment provides a cloud management platform 90, which includes:

[0205] Module 901 is used for:

[0206] Obtain at least one reference image from user input, whereby the at least one reference image is used to describe the normal state of the target object;

[0207] Obtain the first image to be inferred from the user input. The first image to be inferred is used to describe the target object.

[0208] The inference module 902 is used to call the AI ​​model to perform inference on at least one reference image and a first image to be inferred, so as to obtain a first inference result about the first image to be inferred. The first inference result is used to indicate whether the first image to be inferred is in a normal state. The first inference result is obtained based on the difference between the first image to be inferred and the normal state of the target object.

[0209] Optionally, the first inference result is used to indicate that the first image to be inferred is not in a normal state, and the first inference result is also used to indicate the area in the first image to be inferred that is not in a normal state.

[0210] Optionally, the AI ​​model includes a feature extraction layer and a feature processing layer, and the inference module 902 is used for:

[0211] The feature extraction layer extracts features from at least one reference image to obtain the first feature corresponding to at least one reference image;

[0212] The first image to be inferred is subjected to feature extraction by the feature extraction layer to obtain the second feature corresponding to the first image to be inferred;

[0213] The first inference result is obtained by reasoning about the first and second features through the feature processing layer.

[0214] Optionally, the acquisition module 901 is used to: acquire the second image to be inferred input by the user;

[0215] Inference module 902 is used for:

[0216] The second image to be inferred is processed by a feature processing layer to extract features and obtain the third feature corresponding to the second image to be inferred.

[0217] The first and third features are inferred through the feature processing layer to obtain the second inference result corresponding to the second image to be inferred.

[0218] Optionally, the inference module 902 is used for:

[0219] Associate at least one reference image with the user;

[0220] Obtain the third image to be inferred from the user input. The third image to be inferred is used to describe the target object.

[0221] If at least one reference image is identified as being associated with the user, an AI model is used to infer from at least one reference image and a third image to be inferred, in order to obtain a third inference result.

[0222] Optionally, the AI ​​model includes a feature extraction layer and a feature processing layer. The feature processing layer includes one or more sub-processing layers. Different sub-processing layers are used to implement different types of inference tasks. The one or more types of inference tasks corresponding to one or more sub-processing layers include one or more of the following: question answering task, classification task, segmentation task, detection task, and regression task.

[0223] The acquisition module 901 is used to: acquire indication information, which indicates the reasoning task of the target type, and the reasoning task of the target type belongs to one or more types of reasoning tasks;

[0224] The inference module 902 is used to: perform inference on at least one reference image and a first image to be inferred through the feature extraction layer in the AI ​​model and the sub-processing layer corresponding to the inference task of the target type, according to the instruction information, so as to obtain a first inference result.

[0225] Optionally, the AI ​​model is obtained by training based on one or more sets of training data pairs. Any set of training data pairs includes first training data, a first label corresponding to the first training data, second training data, and a second label corresponding to the second training data. The first label is used to indicate that the preset object in the first training data is in a normal state, and the second label is used to indicate that the preset object in the second training data is not in a normal state.

[0226] Both the acquisition module and the inference module can be implemented in software or hardware. For example, the implementation of the acquisition module will be described below. Similarly, the implementation of the inference module can refer to the implementation of the acquisition module.

[0227] As an example of a software functional unit, a module can include code running on a computing instance. A computing instance can include at least one of a physical host (computing device), a virtual machine, or a container. Furthermore, the aforementioned computing instance can be one or more. For example, a module can include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code can be distributed within the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code can be distributed within the same availability zone (AZ) or in different AZs, each AZ comprising one or more geographically proximate data centers. Typically, a region can include multiple AZs.

[0228] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.

[0229] As an example of a hardware functional unit, an acquisition module may include at least one computing device, such as a server. Alternatively, the acquisition module may be implemented using a central processing unit (CPU), an application-specific integrated circuit (ASIC), or a programmable logic device (PLD). The aforementioned PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), a data processing unit (DPU), a neural network processing unit (NPU), a system-on-chip (SoC), an offload card, an accelerator card, or any combination thereof.

[0230] The acquisition module includes multiple computing devices that can be distributed within the same region or in different regions. Similarly, the acquisition module can be distributed within the same Availability Zone (AZ) or in different AZs. Likewise, the acquisition module can be distributed within the same VPC or across multiple VPCs. These computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, GALs, DPUs, NPUs, SoCs, offloading cards, and accelerator cards.

[0231] It should be noted that, in other embodiments, the acquisition module can be used to execute any step in the data processing method, and the inference module can be used to execute any step in the data processing method. The steps that the acquisition module and the inference module are responsible for implementing can be specified as needed. By implementing different steps in the data processing method through the acquisition module and the inference module, all functions of the cloud management platform can be realized.

[0232] This application also provides a computing device 100. As shown in FIG10, the computing device 100 includes a bus 102, a processor 104, a memory 106, and a communication interface 108. The processor 104, the memory 106, and the communication interface 108 communicate with each other via the bus 102. The computing device 100 may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 100.

[0233] Bus 102 can be a Peripheral Component Interconnect Express (PCIe) bus, an Extended Industry Standard Architecture (EISA) bus, a Unified Bus (Ubus or UB), a Compute Express Link (CXL), a Cache Coherent Interconnect for Accelerators (CCIX), etc. The Unified Bus is also known as the Lingqu Bus. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, only one line is used in Figure 10, but this does not imply that there is only one bus or one type of bus. Bus 102 can include pathways for transmitting information between various components of computing device 100 (e.g., memory 106, processor 104, communication interface 108).

[0234] The processor 104 may include any one or more of the following computing devices: central processing unit (CPU), graphics processing unit (GPU), microprocessor (MP) or digital signal processor (DSP), ASIC, FPGA, CPLD, NPU, SoC, offload card, accelerator card, etc.

[0235] Memory 106 may include volatile memory, such as random access memory (RAM). Memory 106 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD) or one or more of these. Furthermore, memory 106 may also be implemented using storage class memory (SCM), phase change memory (PCM), or other types of storage media.

[0236] It is worth noting that the same type of storage medium can be configured in the same computing device to realize the function of memory 106, or two or more types of storage media can be configured to realize the function of memory 106. This application does not limit this.

[0237] The memory 106 stores executable program code, and the processor 104 executes the executable program code to implement the functions of the aforementioned acquisition module and inference module, thereby realizing the data processing method applied to the computing device cluster in the above embodiments. That is, the memory 106 stores instructions for executing the data processing method applied to the computing device cluster in the above embodiments.

[0238] The communication interface 108 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 100 and other devices or communication networks.

[0239] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0240] As shown in Figure 11, the computing device cluster includes at least one computing device 100. The memory 106 of one or more computing devices 100 in the computing device cluster may store the same instructions for performing data processing methods.

[0241] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing data processing methods. In other words, a combination of one or more computing devices 100 can jointly execute instructions for executing data processing methods.

[0242] It should be noted that the memory 106 in different computing devices 100 within the computing device cluster can store different instructions, each used to execute a portion of the data processing method's functions. That is, the instructions stored in the memory 106 of different computing devices 100 can implement the functions of one or more modules within the acquisition module and the inference module.

[0243] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 12 illustrates one possible implementation. As shown in Figure 12, computing devices 100A and 100B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this type of possible implementation, the memory 106 in computing device 100A can store instructions for executing the functions of the acquisition module. Simultaneously, the memory 106 in computing device 100B can store instructions for executing the functions of the inference module. Alternatively, the memory 106 in computing device 100A can store instructions for executing some functions of the inference module. Simultaneously, the memory 106 in computing device 100B can store instructions for executing another part of the functions of the inference module.

[0244] It should be understood that the functions of computing device 100A shown in Figure 12 can also be performed by multiple computing devices 100. Similarly, the functions of computing device 100B can also be performed by multiple computing devices 100.

[0245] This application also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similarly referred to the connection method of the computing device clusters in Figures 11 and 12. The difference is that the memory 106 of one or more computing devices 100 in this computing device cluster can store the same instructions for executing data processing methods.

[0246] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing data processing methods. In other words, a combination of one or more computing devices 100 can jointly execute instructions for executing data processing methods.

[0247] It should be noted that the memory 106 in different computing devices 100 within the computing device cluster can store different instructions for executing parts of the data processing method. That is, the instructions stored in the memory 106 of different computing devices 100 can implement the functions of one or more modules in the acquisition module and the inference module.

[0248] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform a data processing method.

[0249] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center that includes one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to perform a data processing method.

[0250] This application also provides a chip system including a processor for implementing the steps performed by the aforementioned computing device cluster. In one possible design, the chip system may further include a memory for storing necessary program instructions and data. This chip system may be composed of chips or may include chips and other discrete devices.

[0251] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0252] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0253] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0254] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0255] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A data processing method, characterized by, The method is applied to a cloud management platform, which manages infrastructure providing cloud services. The infrastructure includes at least one cloud data center, and at least one server in the at least one cloud data center has an artificial intelligence (AI) model deployed on it. The method includes: The cloud management platform acquires at least one reference image input by the user, and the at least one reference image is used to describe the normal state of the target object; The cloud management platform acquires the first image to be inferred input by the user, and the first image to be inferred is used to describe the target object; The cloud management platform invokes the AI ​​model to perform reasoning on the at least one reference image and the first image to be reasoned, so as to obtain a first reasoning result about the first image to be reasoned. The first reasoning result is used to indicate whether the first image to be reasoned is in a normal state. The first reasoning result is obtained based on the difference between the first image to be reasoned and the normal state of the target object.

2. The method of claim 1, wherein, The first inference result is used to indicate that the first image to be inferred is not in a normal state, and the first inference result is also used to indicate the area in the first image to be inferred that is not in a normal state.

3. The method according to claim 1 or 2, characterized in that, The AI ​​model includes a feature extraction layer and a feature processing layer, and after acquiring at least one reference image input by the user, it further includes: The cloud management platform extracts features from the at least one reference image through the feature extraction layer to obtain the first feature corresponding to the at least one reference image; The cloud management platform invokes the AI ​​model to perform reasoning on the at least one reference image and the first image to be reasoned, to obtain a first reasoning result regarding the first image to be reasoned, including: The cloud management platform extracts features from the first image to be inferred through the feature extraction layer to obtain the second feature corresponding to the first image to be inferred; The cloud management platform performs reasoning on the first feature and the second feature through the feature processing layer to obtain the first reasoning result.

4. The method of claim 3, wherein, The method further includes: The cloud management platform acquires the second image to be inferred input by the user; The cloud management platform extracts features from the second image to be inferred through the feature processing layer to obtain the third feature corresponding to the second image to be inferred. The cloud management platform performs reasoning on the first feature and the third feature through the feature processing layer to obtain the second reasoning result corresponding to the second image to be reasoned.

5. The method according to any one of claims 1 to 4, characterized in that, After obtaining at least one reference image from user input, the method further includes: The cloud management platform associates the at least one reference image with the user; The method further includes: The cloud management platform acquires the third image to be inferred input by the user, and the third image to be inferred is used to describe the target object; When the cloud management platform determines that the at least one reference image is associated with the user, it uses the AI ​​model to infer the at least one reference image and the third image to be inferred, in order to obtain a third inference result.

6. The method according to any one of claims 1 to 5, characterized in that, The AI ​​model includes a feature extraction layer and a feature processing layer. The feature processing layer includes one or more sub-processing layers. Different sub-processing layers are used to implement different types of reasoning tasks. The one or more types of reasoning tasks corresponding to the one or more sub-processing layers include one or more of the following: question answering tasks, classification tasks, segmentation tasks, detection tasks, and regression tasks. The method further includes: The cloud management platform obtains instruction information, which is used to indicate a reasoning task of a target type, wherein the reasoning task of the target type belongs to one or more types of reasoning tasks. The cloud management platform invokes the AI ​​model to perform reasoning on the at least one reference image and the first image to be reasoned, to obtain a first reasoning result regarding the first image to be reasoned, including: According to the instruction information, the cloud management platform performs reasoning on the at least one reference image and the first image to be reasoned through the feature extraction layer in the AI ​​model and the sub-processing layer corresponding to the reasoning task of the target type, so as to obtain the first reasoning result.

7. The method according to any one of claims 1 to 6, characterized in that, The AI ​​model is obtained by training based on one or more sets of training data pairs. Any set of training data pairs includes first training data, a first label corresponding to the first training data, second training data, and a second label corresponding to the second training data. The first label is used to indicate that a preset object in the first training data is in a normal state, and the second label is used to indicate that the preset object in the second training data is not in a normal state.

8. A cloud management platform, characterized by, The cloud management platform is used to manage the infrastructure that provides cloud services. The infrastructure includes at least one cloud data center, and at least one server in the at least one cloud data center has an artificial intelligence (AI) model deployed on it. The cloud management platform includes: The acquisition module is used for: Obtain at least one reference image input by the user, the at least one reference image being used to describe the normal state of the target object; Obtain the first image to be inferred from the user input, the first image to be inferred being used to describe the target object; The inference module is used to call the AI ​​model to perform inference on the at least one reference image and the first image to be inferred, so as to obtain a first inference result about the first image to be inferred. The first inference result is used to indicate whether the first image to be inferred is in a normal state. The first inference result is obtained based on the difference between the first image to be inferred and the normal state of the target object.

9. The cloud management platform of claim 8, wherein, The first inference result is used to indicate that the first image to be inferred is not in a normal state, and the first inference result is also used to indicate the area in the first image to be inferred that is not in a normal state.

10. The cloud management platform of claim 8 or 9, wherein, The AI ​​model includes a feature extraction layer and a feature processing layer, and the inference module is used for: The feature extraction layer is used to extract features from the at least one reference image to obtain the first feature corresponding to the at least one reference image; The feature extraction layer is used to extract features from the first image to be inferred, thereby obtaining the second feature corresponding to the first image to be inferred. The first inference result is obtained by reasoning about the first feature and the second feature through the feature processing layer.

11. The cloud management platform according to claim 10, characterized in that, The acquisition module is used to: acquire the second image to be inferred input by the user; The reasoning module is used for: The feature processing layer extracts features from the second image to be inferred, thereby obtaining the third feature corresponding to the second image to be inferred. The feature processing layer performs reasoning on the first feature and the third feature to obtain a second reasoning result corresponding to the second image to be reasoned.

12. The cloud management platform according to any one of claims 8-11, characterized in that, The reasoning module is used for: Associate the at least one reference image with the user; Obtain the third image to be inferred input by the user, the third image to be inferred being used to describe the target object; If it is determined that the at least one reference image is associated with the user, the AI ​​model is used to infer the at least one reference image and the third image to be inferred to obtain a third inference result.

13. The cloud management platform according to any one of claims 8-12, characterized in that, The AI ​​model includes a feature extraction layer and a feature processing layer. The feature processing layer includes one or more sub-processing layers. Different sub-processing layers are used to implement different types of reasoning tasks. The one or more types of reasoning tasks corresponding to the one or more sub-processing layers include one or more of the following: question answering tasks, classification tasks, segmentation tasks, detection tasks, and regression tasks. The acquisition module is used to: acquire indication information, the indication information being used to indicate a reasoning task of a target type, the reasoning task of the target type belonging to one or more types of reasoning tasks; The inference module is used to: perform inference on the at least one reference image and the first image to be inferred through the feature extraction layer in the AI ​​model and the sub-processing layer corresponding to the inference task of the target type, according to the instruction information, so as to obtain the first inference result.

14. The cloud management platform according to any one of claims 8-13, characterized in that, The AI ​​model is obtained by training based on one or more sets of training data pairs. Any set of training data pairs includes first training data, a first label corresponding to the first training data, second training data, and a second label corresponding to the second training data. The first label is used to indicate that a preset object in the first training data is in a normal state, and the second label is used to indicate that the preset object in the second training data is not in a normal state.

15. A computing device cluster, characterized in that, It includes at least one computing device, said at least one computing device including a processor and a memory; The processor is configured to execute instructions stored in the memory to cause the computing device cluster to perform the method as described in any one of claims 1-7.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when run on a processor, causes the processor to perform the method as described in any one of claims 1-7.

17. A computer program product containing instructions, characterized in that, When the instructions are executed by the processor, the method described in any one of claims 1-7 is implemented.