Method and system for detecting an object in an image
Patent Information
- Application Number
- PCT/EP2026/057759
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-27
- Filing Date
- 2026-03-19
- Publication Date
- 2026-10-01
Smart Images

Figure EP2026057759_01102026_PF_FP_ABST
Abstract
Description
[0001] Method and system for recognizing an object in an image
[0002] The invention relates to a computer-implemented method and a system for recognizing an object in a digital image.
[0003] The invention further relates to a computer program product.
[0004] The development of a foundation model (FM) for visual quality inspection in industry offers several key functions that significantly improve operational efficiency and quality control processes.
[0005] One of the key features is so-called "zero-shot" learning, which allows the model to detect and classify errors or anomalies without having to explicitly train every possible variation.
[0006] This means that the model can effectively generalize to identify new or previously invisible errors based on its understanding of the underlying patterns and features, thereby reducing the need for extensive labeled training data.
[0007] By using advanced machine learning techniques such as zero-shot learning, the basic model can extrapolate from existing knowledge to accurately detect and classify errors even in scenarios where labeled examples are limited or unavailable.
[0008] This function not only streamlines the model training process, but also enables rapid adaptation to changing quality control requirements and dynamic production environments.
[0009] Furthermore, the model's ability to generalize across different product types, materials, or manufacturing processes further improves its utility and applicability in various industrial environments.
[0010] Furthermore, the basic model's ability to learn using zero-shot technology and to generalize significantly reduces the labeling effort required for training.
[0011] Traditionally, labeling large datasets for supervised learning can be labor-intensive, time-consuming, and costly.
[0012] This reduction in annotation effort not only accelerates the model development lifecycle but also minimizes the workload for human annotators, allowing them to focus on complex or fine-tuned quality control tasks. Training a base model in industry is challenging due to the limited and fragmented data availability across different departments within a company.
[0013] Unlike models trained on internet data, industrial datasets are proprietary, sensitive, and subject to strict privacy regulations, leading to data silos and interoperability problems.
[0014] The heterogeneity of data formats, quality, and labeling standards complicates data integration and necessitates extensive preprocessing and harmonization. The scarcity of labeled data further exacerbates these challenges and requires costly and time-consuming manual annotation work.
[0015] Addressing these challenges requires collaboration, coordination, and robust data governance frameworks to ensure effective data exchange and compliance with legal requirements.
[0016] A data catalog is a central repository in which metadata and information about an organization's data holdings are organized.
[0017] It provides a comprehensive overview of available data resources, including descriptions, metadata, and data quality metrics.
[0018] Access policies and permissions are often included to specify who can access or modify the data, along with relevant security restrictions. Usage statistics track data usage and provide insights into the frequency of access to datasets and the purposes for which they are accessed.
[0019] Data ownership information identifies the individuals or teams responsible for managing and maintaining each record within the catalog.
[0020] A data catalog is particularly important in industry for training fundamental industrial models for several reasons.
[0021] The problem with existing data catalogs regarding the findability of suitable data sets lies in their inability to effectively filter and order data sets according to relevance and quality.
[0022] Users may struggle to find the most relevant data records amidst a large number of entries, leading to inefficiencies in the data discovery process. A typical example of the state of the art in this context is a concept using so-called "crawlers" to discover data sources across different platforms, which is based on retrieving data from multiple data sources and then populating its own catalog with relevant metadata.
[0023] However, the well-known AWS "Glue" data catalog and its corresponding crawling capabilities pose a significant drawback for data privacy, as sensitive metadata may be exposed if security measures are inadequate. Automated crawling processes can inadvertently uncover connections between data records, thereby risking the disclosure of confidential information.
[0024] Furthermore, inaccurate categorization or labeling of metadata within the catalog can lead to data breaches or compliance issues.
[0025] Publication US 2025 / 076865 A1 describes a federated machine learning method in which an initially trained machine learning (ML) model is provided by a central model server to multiple clients as their respective local ML models. The initially trained ML model is configured to identify defect features in scanning electron microscopy images. The method also includes the central model server receiving information about an updated local ML model from at least one client. Based on this information, an updated global ML model is determined.
[0026] It is therefore an object of the invention to provide a safe method for visual quality inspection while maintaining privacy and integrity, as well as taking into account data protection and eavesdropping security, which improves the model accuracy and the robustness and adaptability of the model, especially also for domain-specific features, even if a single client of a system only has a small data set for corresponding model training, but further training data sets are available in the system.
[0027] The problem according to the invention is solved by a computer-implemented method for recognizing an object in a digital image by a client-server system comprising a server and at least two clients, wherein the following steps are performed:
[0028] a) Providing a set of local training images to one of at least two clients, each local training image from the set comprising a representation of the object, and the representation containing a valid or invalid representation of the object, which is identified as such; b) Generating and training a local embedding model based on artificial intelligence for a feature of the object by the respective client from the set of local training images using an encoder provided to the respective client in the form of a neural network with input nodes, intermediate layers, and output nodes, wherein the number of output nodes is less than the number of input nodes, and the embedding model is formed by the output nodes of the encoder.
[0029] c) Transmitting the respective local embedding model from each client to the server, and aggregating the respective local embedding models into a global embedding model by the server, and distributing the global embedding model to at least one of the at least two clients,
[0030] d) Capturing a current image of a manufactured product, which includes a representation of the object, by a respective client,
[0031] e) Performing a feature analysis by applying the respective local embedding model, and recognizing the object in the current image if a corresponding feature of the object is present, by the respective client.
[0032] The process uses federated learning (FL) techniques to aggregate the embedding of representations learned from different image datasets across multiple quality control areas, while simultaneously maintaining data protection and confidentiality, and thus system security.
[0033] This ensures privacy, as the data remains local with each data owner to guarantee confidentiality and data protection, which improves system security and integrity.
[0034] Furthermore, domain-specific embeddings are provided, with the embedding models being tailored to the specific area of quality assurance of each data owner and effectively capturing relevant visual features.
[0035] Furthermore, collaborative learning is supported, as the data owners contribute to a collective model without exchanging raw image data.
[0036] Additionally, generalization can be achieved in which aggregated embeddings can be generalized across different quality control scenarios, thereby improving the robustness and adaptability of the model.
[0037] Furthermore, the following steps are performed:
[0038] i) Providing a set of global training images to the server, wherein each global training image from the set comprises a representation of the object, and the representation includes a valid or an invalid representation of the object, which is marked as such,
[0039] ii) Applying the global embedding model to the set of global training images, iii) Determining the similarity between the global embedding model and the respective local embedding models using a similarity function, and identifying at least one local embedding model that lies within a predefined range of similarity values.
[0040] iv) Retrieving the respective set of local training images based on the previously determined at least one local embedding model of its respective client,
[0041] v) Training a provided basic model based on artificial intelligence with at least one previously retrieved respective set of local training images, vi) Capturing a current image of a manufactured product, which includes a representation of the object,
[0042] vii) Recognizing the object in the current image by applying the basic model.
[0043] In a further development of the invention, it is provided that the basic model is a “Residual Neural Network”, RNN, a “deep Convolutional Neural Network”, DNN, according to the standard of the “Visual Geometry Group”, VGG, or an “EfficientNet”.
[0044] This ensures that a basic model particularly advantageous for object recognition is used.
[0045] The standardization body "Visual Geometry Group" has defined an architecture with a deep convolutional neural network (CNN) with multiple layers, also called deep neural networks (DNN), to increase the depth of CNNs and thus improve model performance, where "deep" refers to the number of layers, and, for example, VGG-16 or VGG-19 consists of 16 or 19 convolutional layers, respectively.
[0046] In order to achieve better results from models based on the principle of machine learning, the architectures used became deeper and deeper over time, and multiple CNN blocks were simply stacked on top of each other in the expectation of achieving better results.
[0047] However, DNNs present the problem of the so-called "vanishing gradient".
[0048] The training of a network occurs during so-called "backpropagation", in which, in short, the error travels through the network from back to front.
[0049] In each layer, the gradient is calculated to determine how much each neuron contributed to the error. However, the closer this process gets to the initial layers, the smaller the gradient can become, resulting in little to no adjustment of neuron weights in the front layers.
[0050] As a result, deep network structures often have a comparatively high error rate, which can manifest itself, for example, in insufficient model accuracy.
[0051] However, the decreasing model accuracy cannot be attributed solely to the "vanishing gradient problem".
[0052] However, this problem can be relatively well managed using so-called "batch normalization layers".
[0053] The fact that deep neural networks (DNNs) have poorer performance may also be due to the initialization of the layers or the optimization function used.
[0054] The fundamental building block of a "Residual Neural Network," or RNN for short, is the so-called residual blocks, which incorporate "skip connections." These ensure that the activation of one layer is combined with the output of a subsequent layer.
[0055] This architecture allows the network to simply skip certain layers, especially if they do not contribute to a better result.
[0056] An RNN is composed of several of these so-called "residual blocks".
[0057] EfficientNet is a CNN architecture and scaling method that uniformly scales all dimensions of depth, width, and resolution using a composite coefficient. Unlike traditional practice, where these factors are scaled arbitrarily, EfficientNet's scaling method uniformly scales network width, depth, and resolution using a set of fixed scaling coefficients.
[0058] For example, if more computing resources need to be used temporarily, the network depth, width, and image size can be increased by constant coefficients, with the coefficients being determined by a small raster search on the original small model.
[0059] EfficientNet uses a composite coefficient to scale network width, depth, and resolution uniformly in principle. The composite scaling method follows the intuition that with a larger input image, the network needs more layers to increase the receptive field and more channels to capture finer-grained patterns on the larger image.
[0060] The set of training images includes at least one training image that is a similar representation of the object, where the similarity is determined using a distance function whose value lies within a predetermined range of values.
[0061] Thus, two images are similar to each other if their features can be transformed into one another, for example through transformations.
[0062] In other words, the similarity, i.e., the degree of similarity between the mapping of the object and a corresponding object model, can alternatively be determined using a loss function whose value lies within a predetermined range of values.
[0063] In a further development of the invention, it is provided that the respective encoder is formed by the encoder of a respective autoencoder.
[0064] The autoencoder includes an encoder and a decoder.
[0065] The problem according to the invention is also solved by a client-server system for recognizing an object in a digital image, comprising a server and at least two clients with respective computing devices, wherein the system is configured to perform the method according to one of the preceding claims.
[0066] The computing devices of the server and the clients each include a processor with memory.
[0067] In a further development of the invention, it is provided that an imaging sensor is also included, which is configured to capture a manufactured product as an object using a digital image.
[0068] The problem according to the invention is also solved by a computer program product comprising commands which, when the program is executed by the client-server system according to the invention, cause it to execute the method / steps of the method according to the invention.
[0069] The following figures show an embodiment of the invention.
[0070] Fig. 1 shows an embodiment of the invention in the form of a flowchart,
[0071] Fig. 2 shows an embodiment of the invention in the form of a system. Fig. 1 shows an embodiment of the invention in the form of a flowchart as a method for recognizing an object in a digital image by a client-server system.
[0072] At least one process step is implemented in a computer.
[0073] The system comprises a server and at least two clients as shown in Fig. 2.
[0074] The following steps are performed:
[0075] a) Providing a set of local training images to one of at least two clients, wherein each local training image from the set comprises a representation of the object, and the representation includes a valid or invalid representation of the object, which is marked as such,
[0076] b) Generating and training a respective local embedding model based on artificial intelligence for a feature of the object, by the respective client from the respective set of local training images using an encoder provided to the respective client in the form of a neural network with input nodes, intermediate layers and output nodes, wherein the number of output nodes is less than the number of input nodes, and the embedding model is formed by the output nodes of the encoder.
[0077] c) Transmitting the respective local embedding model from each client to the server, and aggregating the respective local embedding models into a global embedding model by the server, and distributing the global embedding model to at least one of the at least two clients,
[0078] d) Capturing a current image of a manufactured product, which includes a representation of the object, by a respective client,
[0079] e) Performing a feature analysis by applying the respective local embedding model, and recognizing the object in the current image if a corresponding feature of the object is present, by the respective client.
[0080] A client can be a so-called "data catalog owner", i.e., a system for visual quality inspection during the manufacture of products.
[0081] A client can be independent of another client in the system insofar as the inspection systems can be administratively assigned to different manufacturers.
[0082] However, the inspection task should be solved reliably and with as little data as possible. The representation can include a valid or invalid representation of the object, such as correct or defective products or anomalies in the product or product environment.
[0083] Learning to embed images can be achieved, for example, by training a Convolutional Neural Network (CNN) with locally available image datasets. A CNN model is trained to extract one or more meaningful features from the images and capture relevant features for quality control tasks. Transfer learning techniques using pre-trained models such as ResNet, VGG, or EfficientNet can be employed to accelerate the training process and improve performance, especially with limited local data. The resulting local embedding model is then aggregated on the server with other local embedding models from other clients using federated learning, for example, by averaging the individual model weights. This ensures data security, confidentiality, and privacy.
[0084] In particular, techniques such as federated averaging or “secure multi-party computation”, or SM PC for short, can be used to aggregate the embeddings.
[0085] The process and subsequent improvements can be applied iteratively to achieve a gradual improvement of the respective models, for example, even through individual clients who nevertheless trigger an update of the global model.
[0086] Model aggregation can be performed using weights, for example:
[0087] AEM = w * LEM + (1 - w) * GEM
[0088] where
[0089] w a chosen weight,
[0090] AEM, the aggregated image embedding model,
[0091] LEM is the local image embedding model, and
[0092] GEM is the global image embedding model.
[0093] This allows for the application of personalization, where the aggregation step comprises a weighted combination of the LEM and the GEM, and the weight w controls the balance between the two models, enabling flexible adjustment of the influence of the local and global models on the federated learning process. In a specific use case, relevant datasets can be identified using a foundation model (FM).
[0094] In this approach, a user, such as a data scientist, aims to train a basic model for quality control tasks.
[0095] The user has access to several sample images that illustrate the tasks of quality control, such as the training of the basic model for the automotive industry for the press shop.
[0096] Using the global embedding model, the user attempts to identify data catalog owners (DCOs) from other clients of the client-server system whose data sets displayed in the embedding area closely match the provided images.
[0097] In this way, the user can collect relevant training data from these clients to train a robust basic model.
[0098] The procedure also includes (shown in dashed lines in the figure) the following steps: i) Providing a respective set of global training images to the server, wherein each global training image from the set comprises a mapping of the object, and the mapping includes a valid or an invalid representation of the object, which is marked as such,
[0099] ii) Applying the global embedding model to the set of global training images, iii) Determining the similarity between the global embedding model and the respective local embedding models using a similarity function, and identifying at least one local embedding model that lies within a predefined range of similarity values.
[0100] iv) Retrieving the respective set of local training images based on the previously determined at least one local embedding model of its respective client,
[0101] v) Training a provided basic model based on artificial intelligence with at least one previously retrieved respective set of local training images, vi) Capturing a current image of a manufactured product, which includes a representation of the object,
[0102] vii) Recognizing the object in the current image by applying the basic model.
[0103] The basic model can be a Residual Neural Network (RNN), a Deep Convolutional Neural Network (DNN) according to the Visual Geometry Group (VGG) standard, or an EfficientNet. The respective encoder can preferably be formed by the encoder of a given autoencoder.
[0104] Each autoencoder has an encoder and a decoder.
[0105] In other words, an embedding-based prediction can be made in step iii) by having the server receive provided images and using the available global embedding model to predict which clients have image data sets that show similarities in the embedding space to the features extracted from the images.
[0106] This prediction can be based on the proximity, i.e., a short mathematical distance, between the embeddings of the input images and the embeddings of datasets held by different DCOs.
[0107] This allows the user to identify those clients whose image data closely matches the input images, as predicted by the global embedding model.
[0108] These clients are considered relevant for providing their local training data to supplement the training of the base model, as their datasets are likely to contain similar quality-related characteristics or patterns.
[0109] Accordingly, in step iv) the user communicates with the identified clients to obtain access to their data sets or specific samples relevant to the quality control task.
[0110] They thus collect labeled training data from these clients, which may contain images of products, errors or anomalies along with corresponding labels.
[0111] Subsequently, in step v), the basic model is trained using the collected local training data, such as a transformer-based model in combination with CNN layers for feature extraction, to perform quality control tasks.
[0112] Then, optionally, a model evaluation and validation can be performed, in which the user evaluates the trained basic model using validation datasets or real test scenarios.
[0113] Once the user is satisfied with the performance, the trained basic model can be applied to automated quality control tasks on their client.
[0114] The base model can be tested for accuracy using relevant clients identified by the global embedding model to ensure the model's robustness and generalizability. Specifically, the user selects clients close to the embedding area whose data was not used in training the base model to validate its performance across different datasets.
[0115] This test phase verifies the model's ability to generalize to invisible data and ensures its effectiveness in detecting errors or anomalies in different product lines or quality control scenarios.
[0116] After successful validation, the basic model can be seamlessly integrated into the production processes, thus contributing to the maintenance of high standards in terms of product quality and efficiency.
[0117] The described procedure allows for visual quality inspection while respecting privacy, as the image data remains local with each client to ensure confidentiality and data protection.
[0118] Furthermore, domain-specific embeddings can be implemented, whereby the embedding models can be tailored to a specific area of each client's quality assurance and can very effectively capture relevant visual features.
[0119] Through "collaborative learning," clients can contribute to a collective model without exchanging raw image data.
[0120] Furthermore, model generalization can be performed, as aggregated embeddings can be generalized across different quality control scenarios, thereby improving the robustness and adaptability of the model.
[0121] Fig. 2 shows an embodiment of the invention in the form of a system.
[0122] The client-server system for recognizing an object in a digital image comprises a server and at least two clients C1-C3, each with its own computing device and connected cameras CAM1-CAM3 for capturing images of manufactured products. The client-server system is configured to execute the inventive method shown in Fig. 1.
[0123] The imaging camera sensor CAM1 is designed to capture a manufactured product as an object using a digital image.
[0124] The corresponding sensor data from sensor CAM1 are further processed by client C1 as image data in the procedure. Reference character list
[0125] S Server
[0126] C1-C3 Client
[0127] CAM 1 - CAM3 imaging sensor
Claims
Patent claims 1. A computer-implemented method for recognizing an object in a digital image by a client-server system, comprising a server and at least two clients, wherein the following steps are performed: a) Providing a set of local training images to one of at least two clients, wherein each local training image from the set comprises a representation of the object, and the representation includes a valid or invalid representation of the object, which is marked as such, b) Generating and training a respective local embedding model based on artificial intelligence for a feature of the object, by the respective client from the respective set of local training images using an encoder provided to the respective client in the form of a neural network with input nodes, intermediate layers and output nodes, wherein the number of output nodes is less than the number of input nodes, and the embedding model is formed by the output nodes of the encoder. c) Transmitting the respective local embedding model from each client to the server, and aggregating the respective local embedding models into a global embedding model by the server, and distributing the global embedding model to at least one of the at least two clients, d) Capturing a current image of a manufactured product, which includes a representation of the object, by a respective client, e) Performing a feature analysis by applying the respective local embedding model, and recognizing the object in the current image if a corresponding feature of the object is present, by the respective client, characterized by the fact that the following steps are also carried out: i) Providing a set of global training images to the server, wherein each global training image from the set comprises a representation of the object, and the representation includes a valid or an invalid representation of the object, which is marked as such, ii) Applying the global embedding model to the set of global training images, iii) Determining the similarity between the global embedding model and the respective local embedding models using a similarity function, and identifying at least one local embedding model that lies within a predefined range of similarity values, iv) Retrieving the respective set of local training images from the respective client based on the previously identified at least one local embedding model. v) Training a provided basic model based on artificial intelligence with at least one previously retrieved respective set of local training images, vi) Capturing a current image of a manufactured product, which includes a representation of the object, vii) Recognizing the object in the current image by applying the basic model.
2. Method according to the preceding claim, wherein the basic model is a “Residual Neural Network”, RNN, a “deep Convolutional Neural Network”, DNN, according to the standard of the “Visual Geometry Group”, VGG, or an “EfficientNet”.
3. Method according to one of the preceding claims, wherein the respective encoder is formed by the encoder of a respective autoencoder.
4. Client-server system for recognizing an object in a digital image, comprising a server and at least two clients with respective computing devices, wherein the system is configured to perform the method according to one of the preceding claims.
5. Client-server system according to the preceding claim, further comprising an imaging sensor configured to capture a manufactured product as an object using a digital image.
6. Computer program product, comprising instructions which, when the program is executed by the client-server system according to one of the preceding system claims, cause the latter to execute the procedure / steps of the procedure according to one of the preceding procedure claims.